Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rhdunn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
rhdunn
2mo ago
Various libraries (e.g. Python's `re` library) support comments and whitespace as an option allowing you to format the regex on multiple lines with commenting to document what each part does. I'm not sure if there are any regex li
62.
▲
by
rhdunn
2mo ago
Unless they are building a podcast application which is predominantly RSS 2.0 based with some extensions from itunes and others to provide additional podcast-specific metadata such as episode and season information.
63.
▲
by
rhdunn
2mo ago
XSLT 3 (via XPath 3.1) has support for maps, arrays, and parsing JSON to/from those or an XML representation. The XML representation is easier to work with in 3. There's a draft version of XSLT 4 that Saxon supports which adds mor
64.
▲
by
rhdunn
2mo ago
C++ has had smart pointers for memory (and other resource) management for a long time now (see e.g. the Windows ATL classes for working with COM objects and resources). There are a number of challenges that make browsers more challenging (e
65.
▲
by
rhdunn
2mo ago
A common writing tool is to use a story grid. You have chapters/similar along the Y axis and title, characters, plot elements, etc. along the X axis. That way you can keep track of what information you are revealing at each story beat&
66.
▲
by
rhdunn
2mo ago
This has happened with other accelerants in various media/fields: 1. easy access to video recording and editing equipment has made it a lot easier to produce videos on sites like YouTube; 2. easy access to audio recording and editing (
67.
▲
by
rhdunn
2mo ago
IIUC, the main problem with the current Li batteries is that the two plates can over time grow material that will 1) degrade performance; and 2) make it more likely to short circuit and catch fire. Similarly with electric car batteries afte
68.
▲
by
rhdunn
3mo ago
In the linked "Kimi-K3 Technical Report [pdf]" paper, section 2.3 (Stable LatentMoE, p6) has the table with those equations on (top of p7, using β_1 for the gate branch and β_2 for the up branch). They talk specifically about the
69.
▲
by
rhdunn
3mo ago
From the paper (page 6 with a comparison to GLU and SwiGLU) they are not using tanh directly (i.e. f(x) = tanh(x)) but: f_gate(b,x) = b * tanh(x / b) * sigmoid(x) f_up(b,x) = b * tanh(x / b) Looking at the graph I wo
70.
▲
by
rhdunn
3mo ago
Maybe they are in the process of uploading the weights and git history and have taken down the holding page/project to not have the "coming soon" in the git history.
71.
▲
by
rhdunn
3mo ago
I was talking about running this on a server, hence my comments re 1xB200. Obviously, the more hardware/VRAM you have the better/faster you can run these large models. But if you are a small/medium sized company you could fea
72.
▲
by
rhdunn
3mo ago
Yes, that's what I was saying w.r.t. expert offloading, i.e. ensuring that the GPU could fit the active parameters not all the parameters.
73.
▲
by
rhdunn
3mo ago
If it is a mixture of experts (MoE) model like the 2.x models, won't this reduce the hardware needed to run the model? The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough
74.
▲
by
rhdunn
3mo ago
It can be useful for checking input token usage before sending it to the model, e.g. preventing calls above a given token bound or grouping requests into batches. It can also be used by the LLMs to provide the input and output token counts
75.
▲
by
rhdunn
3mo ago
Firefox has had profiles for a long time (via about:profiles and a command-line argument). Unfortunately, the new profiles are not compatible with the old ones and cannot be migrated.
76.
▲
by
rhdunn
3mo ago
1. https://www.uea.ac.uk/about/news/article/fresh-evidence-of-c... -- (2023) Fresh evidence of ChatGPTs political bias revealed by comprehensive new study 2. https://www.ox.ac.uk/news/202
77.
▲
by
rhdunn
3mo ago
I think you're right with the current LLM/transformer architecture. There are several factors that affect model size: - The number of token values supported by the model ("n_vocab"). - The number of parameters/featu
78.
▲
by
rhdunn
3mo ago
Does anyone know if they intend on releasing open source/weights variants for 3.8 or whether 3.6 was the last model they are/were doing that for?
79.
▲
by
rhdunn
3mo ago
A parachuting flamingo? An aardvark driving a bus? It should be easy to randomize the animal and the mode of transport (or vary it with animal playing a sport) to create images not in the training data.
80.
▲
by
rhdunn
3mo ago
Possibly. Telephone (电话) is electricity/electronic (电) + talking/speech (话). In Japanese there's the Japanese possessive ('no') which can also be a modifier/qualifier in text like 男の子 (boy, literally "man
81.
▲
by
rhdunn
3mo ago
Have you tried the https://huggingface.co/LatitudeGames models? They are used by the https://play.aidungeon.com website, but can also be downloaded and used with llama-server in conjunction with something like S
82.
▲
by
rhdunn
3mo ago
It's not goblins, it's honest!
83.
▲
by
rhdunn
3mo ago
The other ingredients would be doing other things: making the pill/drug easer to swallow/consume, extending shelf-life, etc. You need enough of the drug for it to be effective, but not too much to overdose or exhibit side-effects.
84.
▲
by
rhdunn
3mo ago
Yes that clip [4] is making the rounds again in light of Mythos, but his stance (and that of others) hasn't changed ([3] is from a new interview). [1] https://memeburn.com/amodei-says-open-source-ai-is-becoming-... [2
85.
▲
by
rhdunn
3mo ago
LLMs grading the answers is relying on the LLM knowing the answer and not just hallucinating it. You also have issues if/when the model refuses to answer, or if it gets stuck in a loop (e.g. if running locally with a heavily quantized
86.
▲
by
rhdunn
3mo ago
I'm not sure if they've fixed this, but older models have a tendency to ignore negation as `no`, `not`, etc. all occur frequently in the training data so are weighted less strongly than the verbs and nouns. The advice I've he
87.
▲
by
rhdunn
3mo ago
This approach is effectively seeding the context with how you want the LLM to behave/operate ("senior reviewer", i.e. the style of the responses you want) and the context/domain in which the LLM is operating in ("SW
88.
▲
by
rhdunn
3mo ago
You're not reading someone else's RP, though. Not with the applications that support this. There are things like the character(s) you interact with, the initial setting, and some background information predefined, but the response
89.
▲
by
rhdunn
3mo ago
Tell that to the LLM RP crowd.
90.
▲
by
rhdunn
3mo ago
A 27B model can fit easily on a 32GB VRAM card (e.g. 5090) or a 32GB computer in RAM at FP8/Q8 (unsloth have 28.6GB Q8 files). For 24GB VRAM cards (e.g. 4090) you can use Q6_K (22.5GB) or Q5_K_M (19.5GB) quants, possibly offloading som
More ›