Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
agcat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
151.
▲
by
agcat
1y ago
Great piece. I agree its easy to pin point VCs generally in situations like these but ambitious ventures comes with risks.
152.
▲
AI Models Benchmarking for Education
(benchmarks.ai-for-education.org)
3 points
by
agcat
1y ago
|
1 comments
153.
▲
by
agcat
1y ago
I came across this new benchmark leaderboard focused on evaluating AI models for education use cases — found it helpful and well-structured
154.
▲
by
agcat
1y ago
Nowadays, I am working on learning robotics from a software engineer's lens on weekends. I got a yahboom one hand robot with Nvidia jetson nano orin gpu. So far i have setup the robot and next i am going to run basic apps on it. I am
155.
▲
by
agcat
1y ago
This is pretty cool!
156.
▲
by
agcat
1y ago
such a timely advice
157.
▲
by
agcat
1y ago
Congrats on the launch!
158.
▲
by
agcat
1y ago
Thanks for sharing. Its a great writeup
159.
▲
by
agcat
2y ago
You can check out this technical deep dive on Serverless GPUs offerings/Pay-as-you-go way. This includes benchmarks around cold-starts, performance consistency, scalability, and cost-effectiveness for models like Llama2 7Bn & Stabl
160.
▲
Qwen2-7B-Instruct with TensorRT-LLM: consistently high tokens/SEC
(inferless.com)
1 points
by
agcat
2y ago
|
1 comments
161.
▲
by
agcat
2y ago
Hey community: In this deep dive, analyzed LLM speed benchmarks, comparing models like Qwen2-7B-Instruct, Gemma-2-9B-it, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, Phi-3-medium-128k-instruct across Libraries like vLLM, TGI, TensorRT-L
162.
▲
by
agcat
2y ago
You can check out this technical deep dive on Serverless GPUs offerings/Pay-as-you-go way. This includes benchmarks around cold-starts, performance consistency, scalability, and cost-effectiveness for models like Llama2 7Bn & Stabl
163.
▲
by
agcat
2y ago
This is true especially if you are deploying custom or fine-tuned models. Infact, for my company i also ran benchmark tests where we tested cold-starts, performance consistency, scalability, and cost-effectiveness for models like Llama2 7Bn
164.
▲
LLM Wrapper Make Deployment with Nvidia Triton Inference Server Easier
(github.com)
1 points
by
agcat
2y ago
|
1 comments
165.
▲
by
agcat
2y ago
An internal hackathon project to help you deploy with Triton easily. You just write your model logic, run a command, and Triton Co-Pilot does the rest. It automatically generates everything you need, uses AI models to configure inputs and o
166.
▲
by
agcat
2y ago
Yes, we made this as part of an internal hackathon project. We are looking for contributors to make this better.
167.
▲
Show HN: Open-source tool that writes Nvidia Triton Inference Glue code for you
(github.com)
8 points
by
agcat
2y ago
|
2 comments
168.
▲
Open Source CLI Tool to Generate Code for Nvidia Triton Deployment
(github.com)
3 points
by
agcat
2y ago
|
1 comments
169.
▲
by
agcat
2y ago
Triton Co-Pilot: A quick way to write glue code to make deploying with NVIDIA Triton Inference Server easier. It's a cool CLI tool that we created as part of an internal team hackathon. Earlier, deploying a model to Triton was very tou
170.
▲
by
agcat
2y ago
Technical Product Marketer at Inferless.com - You will be taking care of all things content marketing, SEO and Social Media. Understand the needs of our ideal customer persona and craft relevant content. It's 50% Content - 50% Distribu
171.
▲
by
agcat
2y ago
This is a good way to do math. But honestly, how many products actually have 100% utilisation. I did some math a few months ago but mostly on the basis of active users, on what would be the % difference if you have 1k to 10K users/mo.
172.
▲
Real-Time Streaming Apps with Nvidia Open Source Triton Inference
(github.com)
3 points
by
agcat
2y ago
|
0 comments
173.
▲
by
agcat
2y ago
Technical Product Marketer at Inferless.com - You will be taking care of all things content marketing, SEO and Social Media. Understand the needs of our ideal customer persona and craft relevant content. It's 50% Content - 50% Distribu
174.
▲
by
agcat
2y ago
This is a great post. Specially the tricks section. I have been organizing Casual Breakfasts for Developers for last 4 months every 2 weeks. And it's been amazing, have learned so many new things.
175.
▲
by
agcat
2y ago
You can leverage Nvidia Triton Inference server to do this. In case you want a managed Infra service look at this cookbook on how you can use Transformers, Inferless to serve multiple LLMs - https://docs.inferless.com/cookbo
176.
▲
Fast Cold-starts for Serverless GPU Inference is becoming a reality
(inferless.com)
1 points
by
agcat
2y ago
|
1 comments
177.
▲
by
agcat
2y ago
One of our customers partnered with us to use Serverless GPUs for production workloads. They saw benefits like: 1. Dynamic Scaling 2. Reduced Cold start Times consistently at scale 3. Were able to go live in less than one day 4. Maintain se
178.
▲
by
agcat
3y ago
Check this out - https://www.inferless.com/learn/the-state-of-serverless-gpus... Btw, I am the founder of Inferless.
179.
▲
LLMs Tokens/Second Benchmark ( Mistral, Llama2, Gemma) – Independent Research
(inferless.com)
2 points
by
agcat
3y ago
|
0 comments
180.
▲
by
agcat
3y ago
You can also find the same if you know the tokens/sec for different input and tokens variation. In case you are interested to see results for speed ( tokens/second) I Ran some tests between LLama2 7Bn, Gemma 7Bn, Mistral 7Bn to co
More ›