Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Joschkabraun
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Tactics for multi-step LLM app experimentation
(docs.parea.ai)
1 points
by
Joschkabraun
2y ago
|
0 comments
2.
▲
A Systematic Workflow to Build Production-Ready LLM Applications
(docs.parea.ai)
3 points
by
Joschkabraun
2y ago
|
0 comments
3.
▲
Tracking Instructor Validation Errors
(python.useinstructor.com)
1 points
by
Joschkabraun
2y ago
|
0 comments
4.
▲
Generate synthetic data for Q&A tasks via instructor in TypeScript
(docs.parea.ai)
2 points
by
Joschkabraun
2y ago
|
0 comments
5.
▲
by
Joschkabraun
2y ago
What do you mean with being a bad model? If the model is really good at tool use, then it will broadly useful as it needs capabilities to generate the tool definition. So, there should be some transferability.
6.
▲
by
Joschkabraun
2y ago
Ah yes. Have you tried out instructor [0] or Guidance [1]? [0]: https://github.com/jxnl/instructor/ [1]: https://github.com/guidance-ai/guidance/tree/main
7.
▲
by
Joschkabraun
2y ago
Interesting. Do you have any benchmarks?
8.
▲
by
Joschkabraun
2y ago
All models got the same prompt fed which was essentially "Question: {question}". And then the API's accept the function call definition
9.
▲
by
Joschkabraun
2y ago
Have you tried the new beta tool use API? In the experiments I ran there were almost no issues parsing the function call response (similar to GPT-3.5-turbo & GPT-4 turbo)
10.
▲
Anthropic's Haiku Beats GPT-4 Turbo in Tool Use
(docs.parea.ai)
48 points
by
Joschkabraun
2y ago
|
14 comments
11.
▲
Ask HN: Who is building on top of OpenAI's Assistants API?
2 points
by
Joschkabraun
3y ago
|
0 comments
12.
▲
Observability and Testing of OpenAI's Assistants API
(docs.parea.ai)
1 points
by
Joschkabraun
3y ago
|
0 comments
13.
▲
LLM evals on labeled data
(docs.parea.ai)
1 points
by
Joschkabraun
3y ago
|
0 comments
14.
▲
Building and Evaluating Evals for Retrieval
(docs.parea.ai)
2 points
by
Joschkabraun
3y ago
|
0 comments
15.
▲
Reproducible LLM Experimentation with DVC
(docs.parea.ai)
1 points
by
Joschkabraun
3y ago
|
0 comments
16.
▲
by
Joschkabraun
3y ago
https://docs.parea.ai/observability/logging_and_tracing isn't open-source for LLM monitoring but the evaluation metrics to assess LLM app quality are: https://github.com/parea-ai/parea-sdk-py
17.
▲
by
Joschkabraun
3y ago
Co-founder of Parea here, thanks for the mention! We offer testing/evaluation in development ([1], [2]), and production ([3]). You can use pre-built evals ([4]) or create your own ([5]). Do logging ([6]) and go from trace to playground
18.
▲
First LLM playground to visualize log-probabilities for OpenAI models
(app.parea.ai)
1 points
by
Joschkabraun
3y ago
|
0 comments
19.
▲
Logprobs in OpenAI API
(platform.openai.com)
3 points
by
Joschkabraun
3y ago
|
1 comments
20.
▲
Evaluation metrics for any kind of LLM app: RAG, chat, summarization, etc.
(docs.parea.ai)
1 points
by
Joschkabraun
3y ago
|
0 comments
21.
▲
Show HN: Reference-free evaluation of LLM-powered chatbots
(github.com)
2 points
by
Joschkabraun
3y ago
|
0 comments
22.
▲
The first playground to experiment with the GPT-4 Vision API
(app.parea.ai)
5 points
by
Joschkabraun
3y ago
|
1 comments
23.
▲
by
Joschkabraun
3y ago
Happy to launch the first LLM/LMM prompt playground which allows experimenting with the GPT-4 Vision API! Curious to hear what you guys are thinking
24.
▲
by
Joschkabraun
3y ago
Reminds me of "Every Model Learned by Gradient Descent Is Approximately a Kernel Machine" by Pedro Domingos: https://arxiv.org/abs/2012.00152
25.
▲
by
Joschkabraun
3y ago
Yes, GPT-4 is in general much more tuned to respect the system message.
26.
▲
How has the ChatGPT model changed from March to June?
(parea.ai)
5 points
by
Joschkabraun
3y ago
|
2 comments
27.
▲
How to Personalize Stable Diffusion for All the Things
(jina.ai)
3 points
by
Joschkabraun
4y ago
|
0 comments