Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
calebkaiser
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
calebkaiser
1y ago
You can learn an incredible amount. I do quite a bit of research as a core part of my job, and LLMs are amazing at helping me find relevant research to help me explore ideas. Something like "I'm thinking of X. Does this make sense
92.
▲
by
calebkaiser
1y ago
The number of software engineers using IDEs like Cursor is staggering. Among knowledge workers in general, ChatGPT is used widely for basically any task that requires writing or researching. This is obviously not going to be true for every
93.
▲
by
calebkaiser
1y ago
I understand the spirit in this line of criticism, but I think it's easy to muddle the timelines and feel as if things "aren't moving," when in fact, the pace of research and improvement is great. For context: - GPT 2 wa
94.
▲
by
calebkaiser
1y ago
It's a play on the popular quote "Give me a child until he is 7 and I will show you the man", attributed to Aristotle
95.
▲
by
calebkaiser
1y ago
I'm a maintainer of Opik, an open source LLM eval/observability framework. If you use something like LiteLLM or OpenRouter to handle the proxying of requests, Opik basically provides an out-of-the-box recording layer via its integ
96.
▲
by
calebkaiser
1y ago
The 1 million users + $200 million in revenue probably had something to do with the valuation.
97.
▲
by
calebkaiser
1y ago
The author, Calvin Rose, is a (I assume pretty great) compiler engineer at NVIDIA who has a personal affinity for LISPs and Lua: https://bakpakin.com/ I don't think the average programming language enthusiast is mainta
98.
▲
by
calebkaiser
2y ago
Anecdotally, there's been no obvious performance hit, but this is something I should test more thoroughly. I'm planning on running a benchmark across a couple of proxies this week--I'll post the results to HN, if anyone is cu
99.
▲
by
calebkaiser
2y ago
I would recommend looking at OpenRouter, if anyone is interested in implementing fallbacks across model providers. I've been using it in several projects, and the ability to swap across models without changing any implementation code&#
100.
▲
Graph neural networks extrapolate out-of-distribution for shortest paths
(arxiv.org)
2 points
by
calebkaiser
2y ago
|
0 comments
101.
▲
by
calebkaiser
2y ago
It sort of depends on which direction you want to go in. If you're interested in deep RL as applied specifically to LLMs, I'd second another commenter's recommendation of Spinning Up from OpenAI. It hasn't been updated f
102.
▲
by
calebkaiser
2y ago
Fair enough. Happy to scrub.
103.
▲
by
calebkaiser
2y ago
<deleted>
104.
▲
by
calebkaiser
2y ago
Yeah, I'm skeptical about the price point of that particular product as well.
105.
▲
by
calebkaiser
2y ago
The question of "will people pay" is answered--OpenAI alone is at something like $4 billion in ARR. There are also smaller players (relatively) with impressive revenue, many of whom are profitable. There are plenty of open questio
106.
▲
by
calebkaiser
2y ago
In the case of DeepSeek-R1, they used a series of heuristic reward functions that were built for different data types. The paper mentions the use of sandboxed environments to execute generated code against a suite of tests, for example, to
107.
▲
G-Eval for LLM Evaluation
(comet.com)
2 points
by
calebkaiser
2y ago
|
0 comments
108.
▲
Let's Build LLM Judges with Structured Generation
(comet.com)
3 points
by
calebkaiser
2y ago
|
0 comments
109.
▲
by
calebkaiser
2y ago
Eh, I think that's just the state of contemporary "influencer"-heavy discourse. The similarity in how people talk about the two fields has more to do with the people doing the talking (they're the same people) than it do
110.
▲
by
calebkaiser
2y ago
I'm a maintainer of Opik, an open source LLM evaluation and observability platform. We only launched a few months ago, but we're growing rapidly: https://github.com/comet-ml/opik
111.
▲
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
(arxiv.org)
2 points
by
calebkaiser
2y ago
|
0 comments
112.
▲
by
calebkaiser
2y ago
Anecdotally, as part of my job, I've taught ML concepts to a lot of people. I don't know if I've ever worked with anyone who was simply too "unintelligent" to grasp things. The bottleneck, as for most things in my e
113.
▲
by
calebkaiser
2y ago
The short answer is that there is nothing magic about these numbers. Having somewhat standard sizes in the different ranges (7B for smaller models, for example) makes comparing the different architecture and training techniques more straigh
114.
▲
by
calebkaiser
2y ago
The decision to go with Java for the backend was because we feel Java is a bit more battle tested than Python for production (dependency management, concurrency, compilation etc). Go was another strong contender, but we felt like it's
115.
▲
by
calebkaiser
2y ago
Fantastic to hear! Opik should work with OpenRouter out of the box, particularly if you are using the OpenAI Python client to interface with OpenRouter. Opik's integration with OpenAI is implemented via their Python library, and so it
116.
▲
by
calebkaiser
2y ago
Great question. First, we have a ton of respect for the work the DeepEval team is doing. That said, we took a fundamentally different approach in building Opik as an open source project. With DeepEval, if you want to log your data or use th
117.
▲
by
calebkaiser
2y ago
Of the two, the only one I've ever personally explored is OpenLLMetry. Extremely cool project. In general, this is one of those areas where the field still needs to "shake out" a bit.
118.
▲
by
calebkaiser
2y ago
Good question! It mostly came down to implementation speed, as well as some uncertainty about performance/overhead. We will be releasing OpenTelemetry compatible ingestion endpoints in the near future, but since Opik has so many featur
119.
▲
by
calebkaiser
2y ago
Awesome, thanks! If you run into any issues, you can open a ticket on the repo or ping me directly at caleb[at]comet.com
120.
▲
Show HN: Opik, an open source LLM evaluation framework
(github.com)
86 points
by
calebkaiser
2y ago
|
15 comments
More ›