Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
netdur
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
netdur
1y ago
just unsloth on colab using A100 and dataset on google drive.
62.
▲
by
netdur
1y ago
the model runs on H200 in ~20s, costing about $2.4/hr. on L4 it’s cheaper at ~$0.3/hr but takes ~85s to finish. overall, H200 ends up cheaper at volume. my scan has a separate issue though: each page has two columns, so text from
63.
▲
by
netdur
1y ago
I’ve tried that too, trying to detect the scan layout to get better OCR, but it didn’t really beat a fine-tuned Qwen 2.5 VLM 7B. I’d say fine-tuning is the way to go
64.
▲
by
netdur
1y ago
I upgraded to 2TB of storage since 15GB just wasn’t cutting it. Two years later, I’m already at a whopping 200GB!
65.
▲
by
netdur
1y ago
Meta is run by people with no regard for ethics, and if that surprises you, that’s on you. Their whole model is just packaging and selling you with whatever tech they can grab. If you’re worried, don’t install Meta apps. I’ve got WhatsApp o
66.
▲
by
netdur
1y ago
You’re conflating editing with rendering, YouTube didn’t overwrite creators's uploads, it applied an ML filter in the streaming/transcode pipeline, the same layer that already resizes, compresses, and tone-maps. That's not &
67.
▲
by
netdur
1y ago
This outrage feels odd, TV has "improved" movies for ages, youtube doing it with machine learning is the same idea, are we really upset because an ear looks a bit clearer?
68.
▲
by
netdur
1y ago
Man, that whole ‘please make an RSS or JSON feed for me’ request reads like Richard Stallman emailing himself a webpage to print out later, let them use Twitter or whatever medium they wants
69.
▲
by
netdur
1y ago
output tokens must be generated in order (autoregressive decoding), inputs don’t have that constraint, so prefill is parallel, with stronger kernels, KV-cache handling, and batching, Claude can outrun Gemini.
70.
▲
by
netdur
1y ago
I love pricing pages, I avoid landing pages or whatever they want to me read and go directly to pricing page to get the meat of what they offer... then I look at price.
71.
▲
by
netdur
1y ago
Mark Zuckberd must very very upset looking at this, I expect him to throw another billion dollars at google engineers
72.
▲
by
netdur
1y ago
He speaks in movies terms, exactly what I say when I watch movie about programming
73.
▲
by
netdur
1y ago
naive look, 2/3 of model, without multi-languages this shiuld be around 1B
74.
▲
by
netdur
1y ago
it's big little planet or small big planet?
75.
▲
by
netdur
1y ago
I have fine-tuned Gemma 3N 2B and it's pretty good, but loads slow on my S23U, once it's loaded though, it works fine Also tried SmolVLM 256M and 500M, they load faster and you can embed them in assets, they work if you know what
76.
▲
by
netdur
1y ago
Working on https://github.com/netdur/llama_cpp_dart it is llama.cpp binding for Dart first then Flutter I am currently working on multimodal support, add vision and working on audio
77.
▲
by
netdur
1y ago
you must be naive to think OpenAI does not train on your data, Altman is infamous for deceiving claims.
78.
▲
by
netdur
2y ago
My understanding is that in multimodal models, both text and image vectors align to the same semantic space, this alignment seems to be the main difference from text-only models."
79.
▲
by
netdur
2y ago
I followed this guide for fine-tuning: https://ai.google.dev/gemini-api/docs/model-tuning Arabic OCR is a mess with historical texts. Take the word الف (alf/thousand) in dates like 1950 - in old documents, th
80.
▲
by
netdur
2y ago
I have documents from the last 50 years that I need to digitalize, millions of them written in old Arabic. The OCR is not accurate due to handwritten documents, so I need to fine-tune a model on around 300k pairs of texts (OCR output and ma
81.
▲
by
netdur
2y ago
I never heard it was bad at twitter! why is that?
82.
▲
by
netdur
2y ago
Yes, there is Standard Arabic and various Arabic dialects. The Middle Eastern dialects are closer to Standard Arabic while still different. In Morocco, Darija appears to be an Arabic dialect but it's actually more of an Amazigh languag
83.
▲
by
netdur
2y ago
seems on par or better than gpt4 mini
84.
▲
by
netdur
2y ago
it can refuse to do the task based on morals, if it think the output will be used to harm, the issue is not straight refuse, it can subtle refuse by producing results "designed" to avoid accomplish what you want to do
85.
▲
by
netdur
2y ago
Didn't DeepSeek's CEO say that Llama is two generations behind, and that's why they didn't use their methods?
86.
▲
by
netdur
2y ago
While BM25 did emerge from earlier work in the 1970s and 1980s (specifically building on the probabilistic ranking principle), I'm curious about your perspective on a few things: What specific modern statistical approaches are you seei
87.
▲
by
netdur
2y ago
I found YOLOS to be faster and better, bot real time but 22k objects under half second
88.
▲
by
netdur
2y ago
TL;DR: Real-time Linux finally merged into mainline after 18+ years. Good for robots, not your desktop. Real-time kernel ELI5: It's like a super punctual friend who always shows up exactly when they say they will, even if it means they
89.
▲
by
netdur
2y ago
Yes, the size is 99% 2 models weights required to run inference offline, there no way around it.
90.
▲
by
netdur
2y ago
We have not decided what to do with it yet. It could be free, paid, or open source. However, the logic code for using semantic search with CLIP-compatible models on Android will be available on GitHub.
More ›