Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
red2awn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
red2awn
11mo ago
I'd argue SSI and Thinking Machines Lab seem to that environment you are thinking about. Industry labs that focuses on research without immediate product requirement.
32.
▲
by
red2awn
11mo ago
The moat of OpenAI is 1. internal knowledge they've built over the last few years building front tier models 2. their talent 3. the ChatGPT brand (go ask a random person on the street, they know ChatGPT but not Claude or Gemini)
33.
▲
by
red2awn
11mo ago
They are not a profitable company at all. They only started monetizing this year and the customer base is not a fan of it at all.
34.
▲
by
red2awn
1y ago
A lot of optimizations in LLMs now are low hanging fruits inspired by techniques in classical computer science. Another one that comes to mind is paged KV caching which is based on memory paging.
35.
▲
by
red2awn
1y ago
Technically it doesn't have to be since that part of the context window would have been in the KV cache and the inference provider could have thrown away the textual input.
36.
▲
GPT-2 implementation in Modular MAX
(github.com)
2 points
by
red2awn
1y ago
|
1 comments
37.
▲
by
red2awn
1y ago
I am learning to write LLM pipelines using the Modular MAX inference framework. As a starting point I got GPT-2 working after reading through "The Illustrated GPT-2", Karpathy's nanoGPT codebase and existing models in the Mod
38.
▲
by
red2awn
1y ago
The blog post is about LLM non-determinism in the context of serving at scale (variable batch size). The page you link is only about run-to-run determinism implicitly assuming a fixed batch size.
39.
▲
by
red2awn
1y ago
Conceptually setting temperature to be >0 doesn't actually introduce any non-determinism. If your sampler is seeded then it will always choose the same next token. Higher temperature only flattens the logit distribution.
40.
▲
by
red2awn
1y ago
I don't think they were intentionally rug pulling the community. More likely they started out thinking they can build a python superset, realised it wasn't a good idea so quickly pivoted to the current design. Making Mojo more Pyt
41.
▲
by
red2awn
1y ago
If it only takes 30k can't you bootstrap and built it yourself? Or even just work on it on the side alongside your day job.
42.
▲
by
red2awn
1y ago
It also didn't take into account a lot of the new models are reasoning models which spits out a lot of output tokens.
43.
▲
by
red2awn
1y ago
It is literally in the article: > "We're excited that in the next couple months we will be able to share our first product, which will include a significant open source component and be useful for researchers and startups devel
44.
▲
by
red2awn
2y ago
> Architecture Advantages: Enhanced ray tracing, Shader Execution Reordering, and DLSS 3 technology for improved efficiency. This jumps right out as written by AI, these features have nothing to do with training LLMs.
45.
▲
by
red2awn
2y ago
"Chat (formerly ChatGPT)"
46.
▲
by
red2awn
2y ago
Better yet, use doc test as featured in Python [1] or Rust [2]. This makes sure your documentation examples are always up-to-date and runnable. [1]: https://docs.python.org/3/library/doctest.html [2]: https:/
47.
▲
by
red2awn
2y ago
prompt: how many "r"s in the word "raspberry"? response: There are 2 "r"s in the word "raspberry".
48.
▲
by
red2awn
2y ago
The helix editor [1] allows you to navigate and select code objects, I use Alt-o all the time which expand the current selection to the parent syntax node. As mentioned by sibling comments, there are also a bunch of diff tools that are synt
49.
▲
by
red2awn
2y ago
Stopped reading after this > The reason I know this is because I’ve built and maintained systems that handle close to 100,000 payments a day. That's 1.16 payments per second.
50.
▲
by
red2awn
2y ago
> How Do We Prevent This From Happening Again? > Software Resiliency and Testing > * Improve Rapid Response Content testing by using testing types such as: > * Local developer testing So no one actually tested the ch
51.
▲
by
red2awn
2y ago
A very good primer to state-space models (from which Mamba is based on) is The Annotated S4 [1]. If you want to dive into the code I wrote a minimal single-file implementation of Mamba-2 here [2]. [1]: https://srush.github.io
52.
▲
Minimal, single file Mamba-2 implementation in PyTorch
(github.com)
3 points
by
red2awn
2y ago
|
0 comments
53.
▲
by
red2awn
2y ago
It is exactly what this is, pleasing the shareholders. From Tim Cook himself: https://youtu.be/pMX2cQdPubk?t=146 > It became clear that people wanted to know our views on generative AI. So we decided to embrace it and ca
54.
▲
by
red2awn
2y ago
The article seems to be based on misinterpreted information: > Apple’s model has extremely low latency (0.6 milliseconds to first token), outperforms similar sized Phi and Gemini models from Microsoft and Google. From [1] it is 0.6 milli
55.
▲
by
red2awn
2y ago
kitty is being rewritten in Rust from C.
56.
▲
by
red2awn
3y ago
The first one I got was holding a knife...
57.
▲
by
red2awn
3y ago
It is a Blockchain project. https://near.org/
58.
▲
by
red2awn
3y ago
Just use the Rectangle app, problem solved. It would be nice to have it built in, but it is not a big deal.
59.
▲
by
red2awn
3y ago
I experimented with different optimizations and ended with 128x speedup. The improvement mainly comes from manual SIMD intrinsics, but you can go a long way just by making the code more auto-vectorization friendly as some other comments hav
60.
▲
by
red2awn
3y ago
Probably cloned from China.
More ›