Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
newhouseb
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
27 ms
·
31.
▲
by
newhouseb
3y ago
I built a similar thing to Grant's work a couple months ago and prototyped what this would look like against OpenAI's APIs [1]. TL;DR is that depending on how confusing your schema is, you might expect up to 5-10x the token usage
32.
▲
by
newhouseb
3y ago
Ah, this makes total sense. I was thinking about FLOPs in the abstract and not about the wall-clock time. Thanks for the explanation.
33.
▲
by
newhouseb
3y ago
Nope, a decoder only transformer is a variant of the original architecture proposed by Google [1]. All variants of GPT that we know about (1 through 3) all roughly use this same architecture which takes only the decoder stack from the origi
34.
▲
by
newhouseb
3y ago
Ah, you're thinking about embeddings which are basically the encoder stack on a traditional transformer architecture. Modern GPT-like models (including Claude), however, drop the encoder and use decoder-only architectures. I could imag
35.
▲
by
newhouseb
3y ago
This doesn't match my mental model (or implemented model in the case of GPT2) of how self-attention works (you need to calculate the residual stream for each individual token, attending to all prior tokens before it). Have a link?
36.
▲
by
newhouseb
3y ago
Naively, yes, but you can cache the bulk of that "rerunning" [1]. That said the (non-flash) attention costs go up with the length of the sequence so perhaps this is just a simpler way to approximate these costs. [1] https:/&
37.
▲
by
newhouseb
3y ago
I wonder why this is? Naively there's no difference between the two from a transformer standpoint. Perhaps it's because under the hood there's additional safety analysis/candidate generate that is resource intensive?
38.
▲
by
newhouseb
3y ago
Ah yeah, I had the outdated 30M in my head.
39.
▲
by
newhouseb
3y ago
Oh absolutely, I'm just imagining what I might think if I was a super conservative director at Google who is accountable for the balance sheet of a large org.
40.
▲
by
newhouseb
3y ago
But Google hasn't disclosed which version of Bard, right? I pop into Bard every once in a while to test its performance, but I never know if I'm getting the best Google has or just what Google can tolerate running cost-wise public
41.
▲
by
newhouseb
3y ago
This tool is really lovely, great work! I'd be curious to see Softmax Linear Units [1] integrated into the possible activation functions since they seem to improve interpretability. PS: I share your curiosity with respect to things lik
42.
▲
by
newhouseb
3y ago
My emerging conception of this is to split this into two separate questions: 1. Is the architecture _capable_, i.e. is it possible for a model with a given shape possible to perform some "reasoning" 2. Is the architecture _trainab
43.
▲
by
newhouseb
3y ago
Small world! Thanks for the link (which I've now skimmed beyond the abstract). What wasn't obvious to me from the abstract is that different attention heads have different penalty strengths, so if some prediction task requires lon
44.
▲
by
newhouseb
3y ago
First - thank you for open sourcing this! It's a real gift to the community to have a model intended for "commercial use" that's actually licensed as such. I'd be very interested to hear about the choice/evalua
45.
▲
by
newhouseb
3y ago
Thank you! ReLM is a great find! I like that it drives the generation itself so that it can explore different beams more intentionally. And to do the JSON Parsing well against enums/unions/oneOf, you really have to support backtra
46.
▲
by
newhouseb
3y ago
I haven't spent time going deep here but my current hypothesis is that interoperability will more or less end up looking like toolformer where the "tools" are just separate LLM runs with task-specific context. So for example:
47.
▲
by
newhouseb
3y ago
Sometimes, but it very much depends on the context (no pun intended). If it's a pure syntax issue, OpenAI models will almost certainly make the right correction. If it's more abstract, like the LLM has hallucinated a property that
48.
▲
by
newhouseb
3y ago
Oh nice! I built a similar system a few weeks ago: https://github.com/newhouseb/clownfish I think the main differentiating factor here is that this is better if you have a simpler JSON schema without enums or oneOf con
49.
▲
by
newhouseb
3y ago
And I adapted Jay's work to Typescript (without the numpy obviously, just raw typescript/javascript): https://github.com/newhouseb/potatogpt
50.
▲
by
newhouseb
3y ago
The code supports both local and OpenAI as a backend, I added OpenAI as a backend because their models are still miles better than anything I can run locally (even 65B LLaMA)
51.
▲
by
newhouseb
3y ago
Nice! I've been wondering similar things about whether you could use this to eek out more intelligence through methods like these, to quote the end of my write up: > Does structured decoding increase the observability of emergent wo
52.
▲
by
newhouseb
3y ago
There are a couple different approaches: - Use multi-shot prompting with something like guardrails to try prompting a commercial model until it works. [1] - Use a local model with a final layer that steers token selection towards syntactica
53.
▲
by
newhouseb
4y ago
I was having a conversation with some friends the other day about what schemes might mitigate some of these risks, something like an anti-safe word, i.e. a "danger word" that someone can be used to remotely validate that a loved o
54.
▲
by
newhouseb
4y ago
There are a couple different approaches: - Rerun the prompt until you get a format that is consistent - Steer the output token selection towards a predefined prompt For the latter, I've built a proof of concept that takes in a JSON sch
55.
▲
by
newhouseb
4y ago
Last weekend I built some tooling that you can integrate with huggingface transformers to force a given model to _only_ output content that validates against a JSON schema [1]. The challenge is that for it to work cost effectively you need
56.
▲
Structural Alignment: Modifying Transformers (Like GPT) to Follow a JSON Schema
(github.com)
2 points
by
newhouseb
4y ago
|
0 comments
57.
▲
by
newhouseb
4y ago
We also use GPT to perform actions in the software I build at work and we hit the same issue of inconsistency which lead me down a long rabbit hole to see if I could force an LLM to only emit grammatically correct output that follows a besp
58.
▲
by
newhouseb
4y ago
Yes, sorry for not making that more obvious. Since we work with US-based governments at various levels at the moment it's simplest to stick with folks authorized to work in the US. This may change in the future as we look at doing work
59.
▲
by
newhouseb
4y ago
AidKit | Remote | Full-time | https://aidkit.org | $130k + equity | Social Impact / GovTech | TypeScript AidKit runs the largest guaranteed income programs in the country. We replace convoluted workflows of glued together s
60.
▲
by
newhouseb
4y ago
You could configure a second Cybiko, hooked up to your computer, to act as a gateway between an untethered Cybiko and the internet. They teased a standalone device to do this (CyWIG) but I don't think that ever actually shipped. At one
More ›