Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
typpo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
typpo
3y ago
> you may literally see your life's work go up in flames. Incidentally, this happened to Lewicki a few years later when Planetary Resources' first satellite blew up on an Antares rocket: https://www.geekwire.com
32.
▲
by
typpo
3y ago
I like seeing how familiar structures appear at cosmological scale. Long ago I created a webgl visualization of the Millennium Run, an early large-scale cosmological simulation: https://www.asterank.com/galaxies/ It w
33.
▲
by
typpo
3y ago
"Extensions" and integration into the rest of the Google ecosystem could be how Bard wins at the end of the day. There are many tasks where I'd prefer an integration with my email/docs over a slightly smarter LLM. Unli
34.
▲
by
typpo
3y ago
The library supports the model-graded factuality prompt used by OpenAI in their own evals. So, you can do automatic grading if you wish (using GPT 4 by default, or your preferred LLM). Example here: https://promptfoo.dev/doc
35.
▲
by
typpo
3y ago
In case anyone's interested in running their own benchmark across many LLMs, I've built a generic harness for this at https://github.com/promptfoo/promptfoo . I encourage people considering LLM applications to
36.
▲
by
typpo
3y ago
Nice to see a fun, creative project like this. The LLM tie-in makes sense because it keeps headlines fresh & relevant to today's news. Thanks for sharing! If anyone else wants a peek behind the curtain, here is the GPT-4 call: htt
37.
▲
How to benchmark Llama2 Uncensored vs. GPT-3.5 on your own inputs
(promptfoo.dev)
16 points
by
typpo
3y ago
|
0 comments
38.
▲
Benchmark Llama 2 vs. GPT on your own data
(promptfoo.dev)
1 points
by
typpo
3y ago
|
0 comments
39.
▲
Show HN: CLI for testing and evaluating LLM prompts and outputs
(github.com)
2 points
by
typpo
3y ago
|
0 comments
40.
▲
by
typpo
3y ago
As far as I can tell, you are the only person in this thread who actually skimmed the paper. Thank you for pointing this out! The API clearly delineates the March and June versions. The paper authors ran tests on different API versions.
41.
▲
by
typpo
3y ago
I'm responsible for multiple LLM apps with hundreds of thousands of DAU total. I have built and am using promptfoo to iterate: https://github.com/promptfoo/promptfoo My workflow is based on testing: start by defi
42.
▲
by
typpo
3y ago
Thanks for mentioning promptfoo. For anyone else who might prefer deterministic, programmatic evaluation of LLM outputs, I've been building this for evaluating prompts and models: https://github.com/typpo/promptfo
43.
▲
by
typpo
3y ago
For sure. I'm dealing with fuzzier stuff, more in the sense of "don't refer to yourself as a chatbot", "this input should trigger X tool", and things of that nature.
44.
▲
by
typpo
3y ago
The tool maintains an LRU cache on disk by default - which means repeat identical requests will be fetched from cache instead of the live API.
45.
▲
by
typpo
3y ago
This is a great summary of why productionizing LLMs is hard. I'm working on a couple LLM products, including one that's in production for >10 million users. The lack of formal tooling for prompt engineering drives me bonkers,
46.
▲
An open-source framework for prompt engineering
(ianww.com)
3 points
by
typpo
3y ago
|
0 comments
47.
▲
by
typpo
3y ago
Are there established best practices for "engineering" prompts systematically, rather than through trial-and-error? Editing prompts is like playing whack-a-mole: once you clear an edge case, a new problem pops up elsewhere. I
48.
▲
by
typpo
3y ago
Very true. As the author of this visualization, I airbrushed the country borders off the two most recent textures. Because it was too obvious that the borders show the Soviet Union :) Professor Scotese was a great partner and instrumental
49.
▲
by
typpo
3y ago
Looks like the playground is mainly for comparison between models, not prompts, and doesn't support templating? Vercel's is similar but not free and open-source. I'm running these tests in bulk, so I prefer to automate with
50.
▲
by
typpo
3y ago
Thanks for the suggestion! I've added a `promptfoo init` command so the initial scaffolding is much easier.
51.
▲
by
typpo
3y ago
Hi HN, I built this because I'm tuning a bunch of prompts and don't have a great way to do this systematically. This CLI tool helps you pick the best prompt and model by allowing you to configure multiple prompts and variables. It
52.
▲
Show HN: Promptfoo – CLI for testing & improving LLM prompt quality
(github.com)
14 points
by
typpo
3y ago
|
5 comments
53.
▲
by
typpo
3y ago
Hi HN, I maintain a chart generation service, QuickChart ( https://github.com/typpo/quickchart ), which renders millions of charts per day. The most consistent pain point for users is that charts require some programmin
54.
▲
Show HN: Text-to-Chart – embeddable natural language charts
(quickchart.io)
2 points
by
typpo
3y ago
|
1 comments
55.
▲
by
typpo
4y ago
Anyone else unable to complete setup? Mine's been stuck on "Analysing codebase" for hours (~30k LoC).
56.
▲
The Circumnavigators (2017)
(qrp-labs.com)
79 points
by
typpo
4y ago
|
34 comments
57.
▲
Show HN: A discord bot that remixes your friends' profile pictures
(ianww.com)
2 points
by
typpo
4y ago
|
0 comments
58.
▲
by
typpo
4y ago
I've been building something similar that handles the dirty business of formatting a large database into a prompt. Additional work that I've found helpful includes: 1. Using embeddings to filter context into the prompt 2. Identi
59.
▲
by
typpo
4y ago
Hi HN, I built this tool because I found it useful (1) as a way to learn, and (2) to start basic data analysis quickly. It uses OpenAI's completion and embeddings API, although I might be able to move it to a cheaper model eventually.
60.
▲
Show HN: Generate SQL Queries from English
(querymuse.com)
5 points
by
typpo
4y ago
|
1 comments
More ›