Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
robertnishihara
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
Canva Built a Modern AI Platform Using Anyscale
(anyscale.com)
2 points
by
robertnishihara
3y ago
|
0 comments
32.
▲
Building RAG-Based LLM Applications for Production
(anyscale.com)
2 points
by
robertnishihara
3y ago
|
0 comments
33.
▲
Fine-tuning LLMs for longer context and better RAG systems
(anyscale.com)
1 points
by
robertnishihara
3y ago
|
0 comments
34.
▲
Two-day hands-on RAG Bootcamp for developers
(twitter.com)
2 points
by
robertnishihara
3y ago
|
0 comments
35.
▲
RAG at Scale: 10x Cheaper Embedding Computations with Anyscale and Pinecone
(anyscale.com)
1 points
by
robertnishihara
3y ago
|
0 comments
36.
▲
Comparing LLM Performance: Introducing the Open Source Leaderboard for LLM APIs
(anyscale.com)
2 points
by
robertnishihara
3y ago
|
0 comments
37.
▲
LLMPerf Leaderboard
(github.com)
5 points
by
robertnishihara
3y ago
|
0 comments
38.
▲
Anyscale Endpoints: JSON Mode and Function Calling Features
(anyscale.com)
2 points
by
robertnishihara
3y ago
|
0 comments
39.
▲
by
robertnishihara
3y ago
We're hosting the model on Anyscale Endpoints. Try it out here [1] [1] https://docs.endpoints.anyscale.com/supported-models/Meta-Ll...
40.
▲
LLM summarization: A case study of human, Llama-2, & GPT-4 summarization quality
(anyscale.com)
1 points
by
robertnishihara
3y ago
|
0 comments
41.
▲
Reproducible Performance Metrics for LLM Inference
(anyscale.com)
2 points
by
robertnishihara
3y ago
|
0 comments
42.
▲
Building Rag-Based LLM Applications for Production
(anyscale.com)
3 points
by
robertnishihara
3y ago
|
0 comments
43.
▲
Anyscale Endpoints: LLM inference and fine-tuning
(docs.endpoints.anyscale.com)
1 points
by
robertnishihara
3y ago
|
0 comments
44.
▲
Anyscale Private Endpoints and Anyscale Endpoints Fine-Tuning
(anyscale.com)
3 points
by
robertnishihara
3y ago
|
0 comments
45.
▲
Loading Llama-2 70B 20x faster with Anyscale Endpoints
(anyscale.com)
4 points
by
robertnishihara
3y ago
|
0 comments
46.
▲
by
robertnishihara
3y ago
I'm a huge fan of Jax. The Jax team is incredibly strong! Just want to share that Ray (an open source project we're developing at Anyscale), can be used to scale Jax (e.g., across TPUs). Some docs from Google on how to do this ht
47.
▲
by
robertnishihara
3y ago
Here is the blog post accompanying the notebook https://www.anyscale.com/blog/a-comprehensive-guide-for-buil...
48.
▲
A Comprehensive Guide for Building Rag-Based LLM Applications
(github.com)
184 points
by
robertnishihara
3y ago
|
48 comments
49.
▲
by
robertnishihara
3y ago
It shouldn't be 100x. We've built an LLM API at Anyscale, and the price comparison works out as follows (per million tokens) - Llama-2-70B: $1 (on Anyscale Endpoints [1]) - GPT-3.5-turbo: $1.50 - $2 (OpenAI [2]) [1] https:/&
50.
▲
by
robertnishihara
3y ago
Thanks for the feedback, we'll improve the landing page! The models (and current prices) right now are - Llama-2-7B ($0.25 / million tokens) - Llama-2-13B ($0.50 / million tokens) - Llama-2-70B ($1 / million tokens) - Co
51.
▲
by
robertnishihara
3y ago
It's amazing to see how rapidly things are moving. You can try out CodeLlama-34B on Anyscale Endpoints (an LLM inference API we're building here at Anyscale for open source LLMs). https://app.endpoints.anyscale.com/
52.
▲
by
robertnishihara
3y ago
If you want to try out Code Llama, you can query it on Anyscale Endpoints (this is an LLM inference API we're working on here at Anyscale). https://app.endpoints.anyscale.com/
53.
▲
Llama 2 is about as factually accurate as GPT-4 for summaries and is 30X cheaper
(anyscale.com)
19 points
by
robertnishihara
3y ago
|
0 comments
54.
▲
by
robertnishihara
3y ago
We've run experiments on datasets ranging from 5K - 100K examples, which gave fantastic results [1]. Some examples - https://huggingface.co/datasets/b-mc2/sql-create-context - https://huggingface.c
55.
▲
by
robertnishihara
3y ago
I think for fine-tuned GPT-3.5 to be competitive with GPT-4 on your use cases (assistance with Angular), you'd have to fine-tune on enough data that it really resembles pre-training more than fine-tuning. And it wouldn't be worth
56.
▲
by
robertnishihara
3y ago
If you want to query the Llama-2 models, you can use Anyscale Endpoints [1]. Note: I work on this :) Llama-2-70B is $1 / million tokens, which is the most cost-efficient on the market that I'm aware of. [1] https://app.
57.
▲
by
robertnishihara
3y ago
I think of fine-tuning as an avenue to significantly reduce LLM inference costs, so I think this is an exciting development. You're right if you compare GPT-3.5-turbo to fine-tuned GPT-3.5-turbo, but if it's anything like fine-t
58.
▲
Nearly all LLMs will be multi-modal
(twitter.com)
7 points
by
robertnishihara
3y ago
|
1 comments
59.
▲
ByteDance Scales Offline Inference with Multi-Modal LLMs to 200 TB Data
(anyscale.com)
7 points
by
robertnishihara
3y ago
|
0 comments
60.
▲
by
robertnishihara
3y ago
> Llama 2 might by some measures be close to GPT 3.5, but it’s nowhere near GPT 4 I think you're right about this, and benchmarks we've run at Anyscale support this conclusion [1]. The caveat there (which I think will be a big
More ›