Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
huac
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
huac
2y ago
It also feels similar to mixture of depths ( https://arxiv.org/abs/2404.02258 ). Being able to apply this post-training is pretty cool though, makes it easier to use across a wider range of setups.
32.
▲
by
huac
2y ago
> Over the nine-year period of South Korean troop commitments to Vietnam, 40% of the country's overall export earnings during this period came from the money combat personnel were paid, making an average of $200 million each year.
33.
▲
by
huac
2y ago
it's difficult to gauge from outside / as a consumer, but what's interesting is rarely where models are at a given point in time, but rather where the model/team will be with similar amounts resources. it may very well s
34.
▲
by
huac
3y ago
Another useful model to compare to would be DAC https://github.com/descriptinc/descript-audio-codec This is the codec that TSAC extended, so it could be a nice comparison to see. I'd also echo Vocos (from sibling
35.
▲
by
huac
3y ago
There's a number of scaled AMD deployments, including Lamini ( https://www.lamini.ai/blog/lamini-amd-paving-the-road-to-gpu... ) specifically for LLM's. There's also a number of HPC configurations, includi
36.
▲
by
huac
3y ago
I concur english speaking rate: ~150 words per minute gpt tokens per english word: ~1.3 tokens per word 1M hours = ~12B tokens I'd fudge that up a bit because it's probably more tokens from non-English but comparatively seems like
37.
▲
by
huac
3y ago
time to first token != tokens per second remember that EU -> US is ~150ms unavoidable latency, for example. then your comparison is local H100 vs Grok + 150ms latency to first token.
38.
▲
by
huac
3y ago
why don't you stream the results?
39.
▲
by
huac
3y ago
> If I ask to deduce the gender of my voice, can it do that? This iteration is not trained to do so. But the general model structure should work, i.e. if you finetune with instruction data to do so. > Training a projection layer makes
40.
▲
by
huac
3y ago
ChatGPT voice takes the cascaded approach - Whisper to transcribe speech to text, then to GPT, then to TTS. We skip the transcription step. Latency: OpenAI's implementation is quite slow - 5+ seconds to get a reply - but even optimized
41.
▲
Show HN: Real-time voice chat with AI, no transcription
(demo.tincans.ai)
33 points
by
huac
3y ago
|
6 comments
42.
▲
by
huac
3y ago
'narrow scenario,' perhaps, but one that also happens to closely match rumors for GPT4's size
43.
▲
by
huac
3y ago
VSCode is a good target for single-player editing, and I can see something like this being helpful, but what does a collaborative experience look like, eg if you have a team of folks all working on the same prompts?
44.
▲
by
huac
3y ago
the famous FB ads paper (from 10 years ago!) combines decision trees with a logistic regression and shows a significant improvement: https://research.facebook.com/publications/practical-lessons... feel free to extend l
45.
▲
by
huac
3y ago
yeah my code needs to use multiprocessing, which does not play nice with tqdm. thanks for the tip about positions though, that helped me search more effectively and came up with two promising comments. unmerged / require some workaroun
46.
▲
by
huac
3y ago
clever! will have to see if this works with tqdm progress bars, has anyone tried that?
47.
▲
by
huac
3y ago
> 30b+ parameter model doing RAG as part of a conversation with voice responses in less than a second, running on Nvidia. I believe that this is doable - my pipeline is generally closer to 400ms without RAG and with Mixtral, with a lot o
48.
▲
Show HN: A real-time speech-language model for $10 of training
(tincans.ai)
5 points
by
huac
3y ago
|
0 comments
49.
▲
by
huac
3y ago
whisper is simply not designed for this, in many ways, and it's impressive engineering to try and overcome its limitations, but I can't help but feel that it is easier to just use an architecture that is designed for the problem.
50.
▲
by
huac
3y ago
their own shortener, e.g. fb.me, presumably
51.
▲
by
huac
3y ago
yes (you skip a decoding step) but also no (when do you start emitting?)
52.
▲
by
huac
3y ago
> See Midjourney's reddit activity after DALLE-3 What stats are you looking at? Looking at https://subredditstats.com/r/midjourney , I see a slower growth curve after the end of July, but still growing and seem
53.
▲
by
huac
3y ago
do you construct an index over the superbit signatures to perform approximate search or do you perform exact search?
54.
▲
by
huac
3y ago
cheaper prices -> more usage -> larger batch sizes / better gpu utilization -> lower cost of service
55.
▲
by
huac
3y ago
curious what the latency impact is to apply the dimensionality reduction... may hint at the specific technique they use
56.
▲
by
huac
3y ago
I don't think it does. And there is a pretty big risk that you end up picking up on some quirk ("bias") of your reward model that doesn't reflect reality -- GPT4 preferring longer answers is one such commonly observed bi
57.
▲
by
huac
3y ago
I did a similar optimization via https://github.com/viterin/vek as the SIMD version. Some somewhat unscientific calculations showed a 10x improvement staying in float32: https://github.com/stillmatic&#x
58.
▲
by
huac
3y ago
My interpretation is that they prohibit using distilled models to compete against OpenAI, i.e. to offer a foundation model as a service. This particular app is a product not a foundation model so seems pretty clearly fine to me. Of course,
59.
▲
by
huac
3y ago
you could make an informed guess by looking at what the highest quality open source model is, looking at the current employer of that model's creator, and what they currently work on there
60.
▲
by
huac
3y ago
bfloat16 mostly matters for training stability, not for inference. M1 is inefficient for training any reasonably sized model and most inference efforts target 4-8 bits, even.
More ›