Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kkielhofner
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
by
kkielhofner
2y ago
I've used it for hybrid search and it works quite well. Overall I'm really happy to see Typesense mentioned here. A lot of the smaller scale RAG projects, etc you see around would be well served by Typesense but it seems to be rel
2.
▲
by
kkielhofner
2y ago
> I'm an engineering manager How are you involved in the hiring process? > Our engineers are fucking morons. And this guy was the dumbest of the bunch. Very indicative of a toxic culture you seem to have been pulled in to and lik
3.
▲
by
kkielhofner
2y ago
They will. - It's a cornerstone of their brand. - R2 users are at least paying for something (storage). - Their network is massively overbuilt to be able to absorb DDoS attacks. - They offer free bandwidth with their CDN - including co
4.
▲
by
kkielhofner
2y ago
ICC, IPP, QAT, etc are definitely an edge. In AI world they have OpenVINO, Intel Neural Compressor, and a slew of other implementations that typically offer dramatic performance improvements. Like we see with AMD trying to compete with Nvid
5.
▲
by
kkielhofner
2y ago
> The US voting machines are just waiting to be hacked, just a matter of when, not if. The US election system is very distributed and fragmented - there is virtually no standardization. Even in the tightest margins for something like Pre
6.
▲
by
kkielhofner
2y ago
> Do you mind sharing why you chose SPLADE-esque sparse embeddings? I can provide what I can provide publicly. The first thing we ever do is develop benchmarks given the uniqueness of the nuclear energy space and our application. In this
7.
▲
by
kkielhofner
2y ago
> word Wood dominated the embedding values, but these were supposed to go into 2 different categories When faced with a similar challenge we developed a custom tokenizer, pretrained BERT base model[0], and finally a SPLADE-esque sparse e
8.
▲
by
kkielhofner
2y ago
> we don't want to hurt performance on other real-world tasks just to do well on MTEB Nice! Fortunately MTEB lets you sort by model parameter size because using 7B parameter LLMs for embeddings is just... Yuck.
9.
▲
by
kkielhofner
2y ago
LLMs have nearly completely sucked the oxygen out of the room when it comes to machine learning or "AI". I'm shocked at the number of startups, etc you see trying to do RAG, etc that basically have no idea what they are, how
10.
▲
by
kkielhofner
2y ago
My startup (Atomic Canyon) developed embedding models for the nuclear energy space[0]. Let's just say that if you think off-the-shelf embedding models are going to work well with this kind of highly specialized content you're goin
11.
▲
by
kkielhofner
2y ago
> they're not a complete replacement for simpler methods like BM25 There are embedding approaches that balance "semantic understanding" with BM25-ish. They're still pretty obscure outside of the information retrieval
12.
▲
by
kkielhofner
2y ago
Jensen has said for years that 30% of their R&D spend is on software. Needless to say as they continue to crush it financially this number continues to completely race past AMD. Turns out people don’t actually want GPUs, they want solut
13.
▲
by
kkielhofner
2y ago
Shouldn't be much of a surprise, this made news back in 2018 when the same was realized with soldiers and secret military bases: https://www.theguardian.com/world/2018/jan/28/fitness-tracki...
14.
▲
by
kkielhofner
2y ago
For reference seven trillion dollars is 25% of US GDP. Yeah, that's um, wild.
15.
▲
by
kkielhofner
2y ago
Living in a climate (Wisconsin) that has extreme highs and lows my understanding is this is typically intended to smooth-out a gas bill (as one example) moving from $10/mo in the summer to $400/mo in the winter. It’s a budgeting t
16.
▲
by
kkielhofner
2y ago
As one example take a peek at /r/LocalLLaMA[0] (I suspect you know). These people are snapping up anything and everything they can get their hands on at a reasonable price. To your point on the P40, it's an eight year old car
17.
▲
by
kkielhofner
2y ago
> If you get wasted on anything, or do anything silly, or act weird, someone will pull out a phone and video you. Anecdotally this seems to be the key impact. With social media almost everyone now has a "brand" and that brand i
18.
▲
by
kkielhofner
2y ago
> That so many people working in software don’t have deep hardware expertise or are not familiar with data centers plays to that hand. I like to remind myself that AWS is 20 years old. That's an entire generation of people from devs
19.
▲
by
kkielhofner
2y ago
TTL isn't universally respected. Consider the following path: Your machine -> Local router -> Configured upstream DNS Server (ISP/CF/Quad8/etc) -> ? -> Authoritative DNS Server Any one of those layers can ove
20.
▲
by
kkielhofner
2y ago
> Dell/EMC says "Hey, here is drive replacement." We do it, 2 hours later, the volume is knocked offline. Apparently, there was mismatch between backplane version, drive version and through some weird edge case, it knocked
21.
▲
by
kkielhofner
2y ago
> Also, anyone in this industry long enough has been around for "Oh, we will just replace that broken piece of hardware" that ended up "WHY IS EVERYTHING ON FIRE?" because versions didn't match up, hardware was r
22.
▲
by
kkielhofner
2y ago
> They negotiate network and power contracts at a scale that exceeds any typical Fortune 500 company. ..and then mark it up. AWS overall has 38% operating margin[0]. Depending on your application this can hit you really hard (cloud egres
23.
▲
by
kkielhofner
2y ago
> With a part time person you will see downtime when a machine fails If a hardware failure causes downtime you're doing it wrong. Additionally, big cloud scaring people from hardware with marketing and FUD has been very effective. M
24.
▲
by
kkielhofner
2y ago
It’s a joke from a famous moment in HN history: https://news.ycombinator.com/item?id=9224
25.
▲
by
kkielhofner
2y ago
> what model can i run on 1TB With 1TB of RAM you can run nearly anything available (405B essentially being the largest ATM). Llama 405B in FP8 precision fits in H100x8 which is 640GB VRAM. Quantization is a very deep and involved well (
26.
▲
by
kkielhofner
2y ago
Have you seen/heard of Abridge[0]? Long story short their secret sauce comes in two main forms: 1. Accurate speech rec, diarization, etc to record a clinician-patient encounter. No notes, no scribes, no "physician staring at Epic
27.
▲
by
kkielhofner
2y ago
Yes Cloudflare and all of that but they’ll do it for free. Then you get to determine gains you may get from caching and other potential optimizations from one of the best eyeball connected providers in the world. Oh plus the ability to fend
28.
▲
by
kkielhofner
2y ago
An often-ignored/forgotten/unknown fact about utilizing LLMs is that you really need to develop your own benchmark for your specific application/use-case. It’s step 1. “This model scores higher on MMLU” or some other off-the-
29.
▲
by
kkielhofner
2y ago
They get paid per verified kill: https://www.sfwmd.gov/our-work/python-program No idea how it actually works but shooting them where they are has to be tough to verify.
30.
▲
by
kkielhofner
2y ago
llama.cpp and others can run purely on CPU[0]. Even production grade serving frameworks like vLLM[1]. There are a variety of other LLM inference implementations that can run on CPU as well. [0] - https://github.com/ggerganov
More ›