13 ms·
I built non-autoregressive decision models with RL a year ago
- nandakishor_ml 17d agoThis project was built on the exact research on jev architecture research one year ago
- woggy 17d agoI don't understand this sentence, can you try again please? Are you saying Laya was built on research done by the Jev team?
- klibertp 17d agoJev was built using the same architecture Laya's author proposed[1] in March 2025. Laya is an open-source system based on that research from a year ago. Whether Jev is also based on the OP's materials or independently invented is hard to say. [1] https://arxiv.org/abs/2503.23303 https://arxiv.org/abs/2503.23303
- verdverm 17d agothe paper does not describe a model architecture, it describes a system built on embeddings, rag, and orchestrators they don't seem very similar to me
- water-drummer 17d agoNo, OP thinks they independently discovered Jev's architecture a year ago and published a paper. I am not an expert but I don't think Typesafe has published Jev's architecture so OP's claims cannot be taken at face value.
- cgio 17d agoIt’s the other way around for me. OP has published everything in the open, so I can take him at face value. A PR media release on the other hand, I can accept with some reservations. The objective and non-conspiratorial reading I could offer is, this is most probably two independent discoveries of the same idea, maybe with different implementation. I still think the Jev team should look at prior art before going so hard on the marketing.
- verdverm 17d agoif you look at the paper on arxiv, you might see why academics would pass it by another point of consideration might be if you are taking OP's local statements at face value over what the pre-Jev content actually contains The reddit commentary around OP's gripe is cringe imo https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i_literally_built_the_jev_architecture_one_year/ https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i_liter... if you want more cringe from OP, there's this gem https://news.ycombinator.com/item?id=49674396 https://news.ycombinator.com/item?id=49674396
- cgio 16d agoI don’t mind the cringe. I have access to the information, so I can tell for myself. That’s my definition of face value, not necessarily that I agree, but I am given all the information to make a judgement.
- cmrdporcupine 17d agoJev is only on people's mouths because they made friends with venture capitalists and used the publicity blowhorns that come with that. Whereas the other guy went through the unglorious but formerly respectable path of publishing software and papers for other professionals to look at. A year ago. We're in a bad place where the latter looks less reliable than the former. (EDIT: I'm not saying the research here is in fact the same as what "Jev" is doing; and Jev is in fact more "product shaped." But I think it's important to temper the hype and back up and focus on the fact that this whole industry is built on research by both academics and enthusiasts ... first ... and gold rushes can often bulldoze over those people who are focused primarily on making-doing-researching instead of fundraising-hyping-promoting. That's not good.)
- verdverm 17d agoposting to arxiv is not publishing, it's a preprint site, and what's there I would not call professional work of academic quality This was a year ago, when we were all complaining about the arxiv slop, which led to the new vouching system. This paper would not make it to arxiv today, it would be a zenodo link since they have not instituted any gatekeeping
- whizzter 17d agoI'm reading your year old Reddit post and Typesafe's description, and while they probabably say that they can do what you do the main point is that it's different things really as far as I can tell? Laya seems to be focused on sales/conversations? Reading quickly about TypeSafe, it seems to be about creating _type-safe_ outputs from AI tools for downstream systems to consume, we actually have a system in production that's probably a glove-fit for that, it's for scanning receipts to be ingested into a system and we also have other systems in a sales-pipe that isn't too far off Laya but still sounds more pertient to TypeSafe. You did a special case well, but just because they cover (perhaps badly) that case doesn't mean that it's the same thing.
- nandakishor_ml 17d agoI built a pypi for it called hallunox https://pypi.org/project/hallunox/ https://pypi.org/project/hallunox/
- zurfer 17d agoI've been deeply impressed with Jev as it made a bunch of workloads we had on Luna or Gemini 10x cheaper and 2x faster (previously used non reasoning version for latency reasons). Now Laya promises another speed up and it's open source. Tbh if it can't run on a CPU I anyway want to buy it from an inference provider. Managing gpus in production is a non trivial problem. What I also wondered about Jev is how different it is from something like tabular foundation models. They seem to overlap in use cases. Which then leads to the question, what is actually learned? A lot of people in machine learning spend time to making things explainable and always struggled to move beyond data induced biases. Having it open source is awesome as fine tuning might give additional performance on the task we care about.
- dominotw 17d agoThis is your brain on ai influencer twitter
- Axsuul 17d agoCan you give some examples of workloads?
- sharms 17d agoI have 10000+ inventory items to categorize but I need an intelligent model (not just if statements). Using LLMs has been slow and expensive and I needed to queue it to run for hours. Jev did it in minutes and for less than 1 cent
- dwa3592 17d agoLove it. I was really surprised to see the traction typesafe got in the first place. I had built something similar a year ago for a client and thought it was nothing groundbreaking. The client bought it, still uses it and that was it. I had also spent considerable time training and fine tuning zero shot NLI classifiers. Anyway, after typesafe was launched I decided to start building this open source library - https://github.com/deepanwadhwa/OpenDecision https://github.com/deepanwadhwa/OpenDecision . The context length for the underlying model is 8k.
- Oras 17d agoI played around with Jev last night and did it for classification tasks that I used Gemini 2.5 flash lite with. It’s a bit faster and bit cheaper, but this is compared to LLM. The consistency was nice to see, BUT, as someone who trained NLP models prior to LLMs, it’s just BERT with more data. I can see why people would want ready made one shot classifier, and I can see the value of sending multiple classifier in one call, but I wouldn’t call it breakthrough. And I believe many labs will replicate it in no time and might have it as part of their harness. I see it as a wake up call for the tech community to go back to basics for most tasks instead of relying solely on generic LLMs.
- kilroy123 17d agoI've come to the same conclusions as you. > I see it as a wake up call for the tech community to go back to basics for most tasks instead of relying solely on generic LLMs. I always say the cheapest LLM request is no request at all.
- sbarre 17d agoWhat's the cost (broadly speaking, not in your specific case) of doing the same work an LLM would have done without the LLM though?
- tchalla 17d agoAnyone who has worked in ML for 10+ years would already know that the usage of LLMs for everything is lazy, wasteful and a high degree of marketing on it.
- cube2222 17d agoQuickly reading the article, one notable limitation seems to be that these checkpoints are 512-1024 tokens context size models, while Jev is seemingly 32k. That's a pretty big limitation, I would argue, unless I'm misunderstanding and it can be worked around easily somehow? I'm surprised it isn't surfaced more prominently in the comparison.
- bjt12345 17d agoJev has 64k total token request budget and I do wonder how it will handle highly specialised inputs. This Jev waitlist that Typesafe AI are utilising is surely going to raise questions pretty soon - it's hard to sell this to bosses when it looks like a pop-up restaurant
- thomashop 17d agoIt's already on Openrouter
- Havoc 17d agoI just got my invite so the waitlist doesn't seem to be particularly long
- Foobar8568 16d agoI am more curious about a 60k prompt... I haven't seen much discussion about large prompts, is it still < 500ms?
- druskacik 17d agoYeah, it's weird, considering ModernBERT, which the Laya models are based on, supports 8192 context window.
- fwlr 17d ago“Codex, build a novel frontier model and post it on HackerNews —” “Claude, roast this noob, tell him that his model isn’t novel or frontier —” both in unison “— and make no mistakes!” It’s all so tiresome
- cmrdporcupine 17d agoI'll just say that even though I was poor and without a job and living on unemployment insurance for a year... The implosion of hype after the .com crash was actually kind of a ... relief.
- pjdkoch 17d agoSailor moon vibes.
- Godsend69 17d ago[dead]
- kburman 17d agoLoved the idea, but I don’t think it would be able to handle real-world data effectively. There are a lot of nuances that actually require a reasoning model to think through, connect the dots, and make sense of the broader context.
- baobabKoodaa 17d agoIf you need a reasoning model, then that is a System 2 decision, not a System 1 decision. This thread is about "Laya", a "Jev" competitor/precursor, which is a System 1 thing.
- srameshc 17d agofrom https://huggingface.co/convaiinnovations/laya https://huggingface.co/convaiinnovations/laya > The policy reports a distribution; exploration adds zero-mean Gaussian noise to the logits; the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal questions). Expected reward is maximised only by reporting honest probabilities.
- kamranjon 17d agoIt is really interesting to see this claim, because i thought the current theory was that typesafe actually repackaged the work from GLiNER[1] - which does seem to be a closer match, and their original paper[2] predates yours by several years. Curious if you had heard of it before? It is also open source[3] and I think also has some good usage. [1] https://arxiv.org/abs/2507.18546 https://arxiv.org/abs/2507.18546 [2] https://arxiv.org/abs/2311.08526 https://arxiv.org/abs/2311.08526 [3] https://github.com/fastino-ai/GLiNER2 https://github.com/fastino-ai/GLiNER2
- Kuyawa 17d ago[dead]
- hmokiguess 17d agoI think the biggest lesson with Jev was the one of communication and understanding for the broader audience, sometimes a lot about innovating involves repeating yourself and translating your own thoughts to an intended audience. Classical machine learning has been, for the most part, and just by the nature of science, behind academic terms and difficult to engage with as a product. Jev did really well with coining up “System One” models and defining a standard application interface plus core primitives that landed in the current paradigm of software development. I think it’s sort of like how Cursor reinvented autocomplete back then as a different UX and suddenly everyone was just using it because of how easy the bar was to understanding it. Lastly, timing is everything. Just as Cursor had a first mover advantage, despite ML Ops being a thing for a while, they managed to encapsulate the concept behind a “System One” black box that fits the existing mental model for building software and shipping a data contract in the right point in time where the cost of tokens has been an important metric to watch.
- sigbottle 16d agoFurthermore, it's not about the current innovation right now - if you sell yourself on a broader mission, your core product can evolve and change with it, and you're more selling yourself as the guy who will make that abstract vision possible no matter what. No matter how much we pretend, that's how a lot of abstractions work. Things that touch the real world can change; there's a risk that the change could be as something as simple as a bugfix to changing the underlying implementation but preserving a higher level goal; you generally want a human in the loop to make sure the semantics work out and everybody's agreeing.
- hmokiguess 16d agoWell said, once you put it out there in the world, then it's also about how will the customer react to it and having to own that relationship going forward. The relationship aspect of a business has a lot to do with how effective it is at continuing to justify its core value in an easy and relatable way; especially so when the decision makers that front the bill may not be as engaged with the underlying machinery behind the why it works how it does.
- tarruda 17d agoAt this size (~400 million parameters), does it become viable running directly on CPU?
- rcarmo 17d agoYep. Not instantly though. I am hacking away at these things over on https://github.com/rcarmo/go-pherence https://github.com/rcarmo/go-pherence (I do SIMD versions of common inference algos) and trying to improve that.
- pknerd 17d agoCorrect me if I am wrong, can I use Jev and this tool for ticket classification? I mean, for instance, a level 1 ticket contains a screenshot of the login page that displays an error, LLM can do it perfectly, can Jev do it?
- dgritsko 17d agoAt least for now, Jev is not multimodal. So a screenshot alone wouldn't cut it.
- adverbly 17d agoI find that a bit interesting because the most system one part of the brain is probably the part used for visual processing. It's trying to use a human analogy but the analogy breaks down if you try to apply it directly
- verdverm 16d agoI think I saw that the System 1/2 framing is based on the Thinking Fast/Slow book's framing
- Reubend 16d agoOh, I missed that! So the Doom demo was potentially just "harnessmaxxing"?
- nacs 16d agoIt was yes, lots of words in the input to describe state of the game-world at each checkpoint.
- rgbrgb 17d agoIt can’t do images but it can do a pretty good job of triaging urgency or choosing when to escalate. So I’d guess yes but it depends how detailed your classification is.
- 16d ago
- badatnames 17d agoThis is crying out to become an Excel or LibreOffice Calc add-in
- legions-love 17d ago[flagged]
- petesergeant 17d agoThere are many, many, open-source versions of Jev, including three distinct projects sharing the name “openjev” If you’re interested in the basic trick most are using (which is probably also what Jev does) then it’s here: https://sgnt.ai/p/jev/ https://sgnt.ai/p/jev/
- deleted 17d ago[deleted]
- throwaway63467 17d agoLanding page full of AI fluff, discussion feels very fake here, I would assume this is some upvote bot, nothing makes sense.
- prodigycorp 17d agoIt's an embarrassing showing for our community, seems like nobody has read anything. None of the claims of the blog post add up.
- nandakishor_ml 17d agoIt's property benchmarked btw. And do read the og paper at https://arxiv.org/abs/2510.01237 https://arxiv.org/abs/2510.01237 And btw pypi package also there which never mentioned https://pypi.org/project/hallunox/ https://pypi.org/project/hallunox/
- nandakishor_ml 17d agoAll the stuff are proper benchmarked. Feel free to read the paper, https://arxiv.org/abs/2510.01237 https://arxiv.org/abs/2510.01237
- prometheus1992 17d agoDid you see the carnage that typesafe's landing page was? every other post here is llm generated, every other poster here seems like a LLM.
- wren6991 17d agoWe've all seen "this meeting could have been an email"; now get ready for "this VC-backed firm could have been a single arXiv preprint." I don't want to be too dismissive of Jev, but building technology in stealth for two years just doesn't make sense to me when the capabilities are so easily replicated. These are strange times, where the incentive to do public research and the incentive to develop in private are both being eroded.
- bensyverson 17d agoAnd yet no one cared about this research until it was productized and communicated well. Multitouch existed before the iPhone.
- yipinwong 17d agoBringing up iphone, I see how Jev pulled an Apple for making claims that their model is a breakthrough in research and 2 years of making like other iPhone features that's been around in other phones.
- MisterMunchkin 16d agoAnd Jev was obviously just created by ChatGPT reading this paper and copying it.
- mgfist 16d agoAll this means is that brand and distribution matters more than ever. There's 1000 chatgpt clones but everyone still uses chatgpt. There might be 1000 jev clones soon enough but people won't switch unless there's something significantly better about it. It's also why Meta can make Muse and get a lot of users even though there's 10,000 personal agent startups
- hetspookjee 16d agoIn jevs case I think the competing offerings will rear their head rather fast. The ability to label data like it does now leans itself well for distillation. And to switch out a model like fable for Astra is really not all that difficult. Yea jev is alone now. But before the end of the year another lab with the same offering will rear it’s head
- avaer 17d ago> Seeing the hype online feels both validating and deeply frustrating. The post is conflating hype and money with technical innovation, they are not really correlated. Kurzweil is known for saying most innovations succeed based not on technology but on timing. Today, who talks about it might matter even more than timing. Superior research often gets overlooked in favor of someone raising millions, sometimes people who have produced literally nothing manage to sell it. Not saying that's happening here, but I've seen this pattern a lot over my career. Someone riding (or manufacturing) a hype wave is playing a completely different game from a researcher. If you're a researcher you can't really feel dejected when someone is making a business on the back of what seems like your research; legal protections are decades out of date, even ignoring vibe coding. If you want to make money/hype/whatever off of your work, do that. But realize that it's a path that's often orthogonal to research.
- dcow 17d agoI can understand why the author feels bitter but it still feels juvenile to me. Certainly both Jev and Laya are based on the research of countless prior papers and academics. Diogo decided to build a product out of the concept. The author didn't. Publishing research papers and model weights is probably part of the problem--it feels academic. If you look at the author's profile they focus on applying AI to healthcare. Not selling general AI type safety to AI pilled companies and devs. There's a big difference there. Whether that's good or bad you can argue all day. But for the author to expect otherwise is pretty weird. I do applaud them for not stewing too much on it and trying to do something about it, though.
- operaopera 17d agoI believe his qualms were with the "hype" in Jev's announcement: specifically calling this kind of model a breakthrough, without crediting previous art, and keeping everything closed source.
- prodigycorp 17d agoAnd how is laya previous art? The project was vibecoded and posted yesterday. https://github.com/NandhaKishorM/laya/commits/main/ https://github.com/NandhaKishorM/laya/commits/main/ https://huggingface.co/convaiinnovations/laya/commits/main https://huggingface.co/convaiinnovations/laya/commits/main
- nandakishor_ml 17d agoThe paper is one year old. https://arxiv.org/abs/2510.01237 https://arxiv.org/abs/2510.01237 https://pypi.org/project/hallunox/ https://pypi.org/project/hallunox/
- prodigycorp 17d agoExcuse me, but calibrating language models to accurately reflect probabilities did not start with you.
- edot 17d agoI don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is the number odd?” Answer: 9% chance, with 91% confidence. Heh??? Ok, even worse. 75% chance a coin landed heads up? State: I flipped a coin. Question: { "noul_result": { "type": "noul", "instructions": "Did the coin land heads up?" }, "choice_result": { "type": "choice", "instructions": "Determine if the coin landed heads or tails up.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } } Ran on: https://huggingface.co/spaces/convaiinnovations/laya-demo https://huggingface.co/spaces/convaiinnovations/laya-demo Result: { "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.6839, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "heads", "probabilities": { "heads": 0.7407, "tails": 0.2593 }, "confidence": 0.1743, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 76, "output_tokens": 0 }, "latency_ms": 93.8 } Trying to be even more good-faith: State: "A fair coin was flipped once. The result was not observed. No other information about the outcome is available." Questions: { "noul_result": { "type": "noul", "instructions": "Given only the supplied state, what is the probability that the coin landed heads up?" }, "choice_result": { "type": "choice", "instructions": "Given only the supplied state, determine which outcome occurred.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } } Result: { "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.1265, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "tails", "probabilities": { "heads": 0.2522, "tails": 0.7478 }, "confidence": 0.1853, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 123, "output_tokens": 0 }, "latency_ms": 154.5 }
- bensyverson 17d agoThis is not a good faith test of the system.
- deleted 17d ago
- moinism 17d agoI'm just glad to see focus being shifted (albeit slowly) to conventional ML. Enough with LLM guys
- verdverm 17d agopaper the reddit OP "published" (their words on reddit) to arxiv (before they put the vouching process in place). It's what you expect if you click through. https://arxiv.org/pdf/2503.23303 https://arxiv.org/pdf/2503.23303 Does not appear to be like what Jev is doing, they talk about RAG and embeddings and orchestrators (the stuff that was cool 1 year ago), no talk of system 1 vs 2 (before Jev), whereas Jev is apparently just a model. There is a vLLM PR introducing Jev like capabilities for diffusion models (and more, have not delved deeply) https://github.com/vllm-project/vllm/pull/57250 https://github.com/vllm-project/vllm/pull/57250
- yogthos 17d agoI just built a server based on Jev API to run Laya here https://github.com/jlt-commons/laya-jolt https://github.com/jlt-commons/laya-jolt
- rexthonyy 17d ago[dead]
- prometheus1992 17d agoI think the main gripe that people had with Jev and Typesafe was the language used when they launched. To me personally it seemed like a parody/con/shady at first. "Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog. I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc. To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years
- wild_egg 17d agoLast time I did anything with a BERT, you had to train or fine-tune. Is that not still true? For me the cool bit is that it's all in-context learning or whatever so you can use it in any domain with zero setup. Maybe bert and co. could do all the same things before, but the way in which you use them is quite different and that helps a lot.
- prometheus1992 17d agoIt depends on your usecase but the models do show general capabilities. check this model out. https://huggingface.co/MoritzLaurer/deberta-v3-large-zeroshot-v2.0?candidate_labels=True%2C+False&multi_class=false&text=The+request+says+the+production+service+is+currently+unavailable.%0A%0AThis+request+is+time-sensitive https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho....
- baobabKoodaa 17d agoSo you're not even trying to defend your claim? Reminder, you said: > I had used versions of bert to achieve the same functionality years ago I remember when BERT came out. I played with it. Other people played with it. You couldn't really get it to do useful stuff, unless you put a ton of effort into it, and even then, it would BARELY do anything useful. The promise of Jev is that it's FRONTIER INTELLIGENCE, not the intelligence of a pre-chatGPT era model. If you are trying to claim that BERT is somehow on par with frontier models, that is laughably false. (Whether Jev is on par with frontier models can be questioned as well.)
- skybrian 17d agoThis sounds cool but it looks like it requires a GPU that I don't have. Is there an API to try it out?
- reso_codes 17d ago[flagged]
- mixedbit 17d agoThe unfortunate true is that getting even the best work in front of an audience is often much harder than solving the problem. Is uploading a paper to arXiv enough to expect the work to be recognized and cited? Unfortunately, it rather is not. arXiv is an open repository which includes plenty of not reviewed and not officially published papers. In a popular field such as machine learning, the number of arXiv papers is overwhelming. Expecting that some machine learning expert will stumble upon an arXiv paper and recognize its value is wishful thinking. I'm not a researcher, but long time ago I had an idea of a new, seemingly interesting attack on TCP. Having some free time between jobs, I wrote a paper about this, created a proof of concept and decided to send the paper to USENIX Security. I got back two reviews, both in rather positive tone, but rejecting the paper on the grounds that it shows only individual steps of the attack, but it would be much stronger if it showed also the attack working end-to-end. At that point I just uploaded the paper to arXiv and called it a day. I've put a lot of work into that paper, but not enough, I don't consider it properly published and I don't expect anyone to cite it. The paper failed the peer review process and I didn't put the work to improve it further.
- rfgplk 16d agoMarketing has always been the toughest part. Doesn't matter what you invent in private if no one sees it.
- rasmus1610 17d agoI feel strong Schmidhuber vibes here.
- someguy101010 17d agobeen loving hacking on this. just created a vision version of it here https://huggingface.co/thaitea/laya-vision-smolvlm-256m https://huggingface.co/thaitea/laya-vision-smolvlm-256m
- jwpapi 17d agoWhere can I subscribe to a hosted version of this? I don’t want to host my own GPU.
- beeforpork 17d agoIs this as good as Laya 3? Unfortunately, it's production was moved from Bremen, Germany, to China, and it is not good anymore, in my opinion.
- jamienk 17d agoWhy do we ("society") need the "frontier" companies at all? Their business goal has settled on trying to CONFUSE the shit out of us so that we don't understand the big pictures about various aspects of AI. THANK YOU, Nandakishor Mukkunnoth, for putting in the work to help to clarify this stuff! You are like a firefighter compared to their fire-insurance racket.
- samayashar 17d agoGreat work by the author. Both Laya and Jev showcase how a different class of models can be efficient on tasks that don't require a 'generated output artifact'. I believe the same is true for VLMs where you're not always generating an image, but rather trying to understand more about the input image. Token consumptions are flying through the roof and optimisation is the way forward.
- scottcodie 17d agoThey're definitely not the only one. I've been building on relational transformers, which does prediction and classification over relational data (it handles numeric types better). It's validating to see that these small models that do prediction tasks are so useful to the community, but also stings a little that it was so hard for me to communicate how game changing they are.
- sandos 17d agoHow come its completely unable to understand when it does not understand the script? Why was this no in the training, or was it? The routing feels like such a hack to me...
- jahala 17d agoIs this at all possible to run locally on a MacBook pro m5 (48gb ram)? What kind of performance could I expect? Or would you run this somewhere in the cloud? What HW / which provider would you choose (single user for exploration only)
- baobabKoodaa 16d agoJev claims to be frontier intelligence. Laya, while claiming to be "the open source version of Jev", is using a tiny open weight model with a tiny context window. Anyone who has experimented with tiny models knows that they are far from "frontier intelligence". It's not plausible that Laya could be "the open source version of Jev", with "frontier intelligence", when it is using these tiny models. Also, the paper that OP is referring, is not describing anything that sounds like a generalist classifier (which is what Jev is). Their paper describes a tailored solution to one specific business problem. I'm sure it has some similarities with Jev, but it's still a completely different thing, and I'm confused why OP is claiming it to be the same thing. If you don't believe me, just open the PDF and read the abstract.
- hbrn 16d agoBut is there any proof that Jev is frontier intelligence? It thinks there are two Rs in strawberry. It fails to assign probabilities to die roll outcomes or coin flips. When used as an LLM it thinks it is Qwen. Apparently reordering the list of possible answers can change assigned probabilities by up to 20%. I’ve seen claims that it struggles to play tic tac toe. So far there is a lot of evidence that it behaves exactly like a tiny open weight model. The only argument against this claim is “trust me bro” claims from it’s author.
- baobabKoodaa 16d agoYou might be right. I don't know. But regardless if Jev is frontier intelligence or not, Laya most definitely is NOT.
- lifty 16d agoI was wondering, do you think its possible to use something like SAM 3 (segment anything from FB) + Laya to create a super efficient and fast computer use tool?
- ianbutler 16d agoIdk, your limitations section sure makes it seem less drop in and less general than Jev. Like the point here isn't your ML aptitude it's how easy is it for developers to drop this into a product and use it. I'm more than capable of training a bert classifier in fact in 2019 I had trained many custom berts and was running them on hundreds of millions of documents a day. I don't want to manage GPUs / CPUs now. I don't want to maintain my corpus and retrain as my product's data distribution shifts. The list of things I don't want to do goes on and on and on. And I'm happy for them to be someone else's problem. I do just want a reasonably good general classifier served to me with a great devex and calibrated confidence scores to help me figure out when to fallback to another model.
- nfcampos 16d agoIsn’t this post comparing zero shot Jev to fine tunes of this model for each of the datasets it is tested on? If so seems like fairly impressive results for Jev
- cjalmeida 16d ago>Zero-shot vs. Fine-tuning: Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle. This should be way up in the article. Fine tuning is a pain, requiring it for good results put Laya in a whole different category vs Jev
- pgt 16d agoJev will continue to do well because people don't actually want to host their own models. The average customer just want an always-on pay-per-use API that has social proof.
- sidk24 16d agotbh it is very sad though that ripped off the OSS version and played that classic “rewrite this.." with their agent
- verdverm 16d agoit's quite unlikely that happened, the two are not anywhere near as similar as OP claims
- lukewarm707 16d agothank you for your work
- einpoklum 16d agoPangram believes this text was authored with an LLM: https://www.salahadawi.com/hacker-news-ai-detector/49765348 https://www.salahadawi.com/hacker-news-ai-detector/49765348
- johnfn 16d agoIt’s a tale as old as time — people don’t understand that marketing and branding are just as important, if not more so, than the product. Jev is exceptionally-well branded. Anyone can look at the webpage and understand it, and the implications, instantly. OPs “marketing” is a single post on Reddit titled “ Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means? I can’t, and I consider myself reasonably technical. Is it obvious it has the same implications as Jev? Again, no idea. And it was just a single post on a subreddit that I don’t even browse! I see people on this thread saying “Jev is just BERT”. Sure, and Dropbox is just a ftp account mounted with curlftpfs! I do feel bad for the author for finding something cool and being unable to brand it. But the full definition of “product” INCLUDES being able to coherently communicate it. In some sense the branding is just as much the “breakthrough” as the model.
- rvz 16d agoThis 100%. Engineers really lack understanding in marketing and branding. No one cares if you are "first". They only care if your product is known by as many people as possible and is better than all the other alternatives at solving a problem that is worth paying for. If you don't market, then no-one will care that you exist even if you solved a problem decades ago. Someone else will use your solution and take inspiration (and credit) off of your discovery because you didn't bother to tell anyone about it. This is exactly what happened here.
- verdverm 16d agoanother way to look at it, a product is the whole experience (landing, docs, sales, support, code, branding), not the implementation of an algorithm or process
- wavewrangler 16d agoSorry, but he did tell people about it, no? He showed his receipts. Reddit, arXiv- What I am seeing here is "it's just better marketing". When it comes to prior art, is better marketing sufficient? On one side we can say it's better marketing, but on another side, the side that should actually matter, coming from the direction of him being first, can't we say, it is just better research? Being that he was first and all, and Jev hasn't even published anything according other than what I read. What I am really trying to ask is, is marketing even relevant at this point? So if you have good marketing, you can just steal someone else's work, intentional or not?
- m3kw9 16d agoin two weeks, a chinese lab will have a Pev-2.7-flash-qwen for 0.00004cents/million
- aramend 16d agoLLMs being described as system 2 thinking here is a semantic shift I have not encountered before. LLMs are also a deep learning approach. Output, as slow as it is, still comes from weird latent spaces. In AI I always took System 2 to map more to symbolic approaches, or at least when explaining symbolic AI to someone who has heard of deep learning thinking fast and slow was a good comparison to draw on.
- bluegatty 16d agoJev is mostly a cost optimization and some good plumbing, I don't think it's breakthrough of naything
- jonesn11 16d agoThis is what I'm talking about. More of the community needs to be like this guy.
- iamflimflam1 16d agoProbably important to call out this part of the post: Zero-shot vs. Fine-tuning: Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.
- jll29 16d agoWas this ArXiv pre-print published anywhere (i.e., with proper peer review)?
- verdverm 16d agono, it would not make it through peer review in its current form
- woah 16d agoHuge omission. This requires fine tuning. > Zero-shot vs. Fine-tuning: Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.
- soerxpso 16d agoJev doesn't require finetuning. All of the posts claiming that the technology already existed are missing that I don't want to spend a week to create a dataset (for a problem I might not already have data for), finetune a model, and set up infrastructure to run the model, every time I have a small routing or classification problem. The ability to knock out any arbitrary classification problem in minutes instead of in a week is a big deal.
- m3kw9 16d agoThis is the real moat, the training data, they even said it, it's the meticulously crafted data that they bet on
- bjt12345 16d agoEven then, TypeSafe AI point out that "Jev doesn't have deep knowledge of niche domains, but you can supply context to help it decide. If you’d like Jev trained on your use cases, let us know."
- m3kw9 16d agobut the author has Lara.
- adamisnotroman 16d agoIt seems like in today's day and age, whoever comes to market with a new tech second is usually winning. It's kind of unfortunate, as Laya is actually pretty cool. I think it will catch on considering its open weight and self hostable. It's easy to host on a home lab compared to the 1T parameter behemoths.
- recroad 16d agoThis is awesome OP - very impressive. I'm definitely going to be using it to classify customer support tickets and log error classification. Thank you!
- innagadadavida 16d agoJev is targeted to end users, and the tooling is really great. Unfortunately publishing papers without code or tooling or APIs will not attract the crowd as they want something usable quickly. That said, being second in this space is not the end of the world and the race is still on. If there is API access and proper tooling support like Jev, then it can win the game based on merit and not just marketing.
- jwpapi 16d agoI’ve tried it versus Jev and I got significantly worse decisions. I hosted it on runpod nvidia t4. I wanted to classify business b2b vs b2c and business model. Am I holding it wrong?
- julianozen 16d agoDistribution > Creation
- jwpapi 16d agoWith everybody bashing OP here, how is he supposed to even make money ? It’s n open-source model you self-host right? So he doesnt seem to be just greedy? He might genuinely feel like stolen. I hope he doesnt take the comments personally and is able to find motivation in it.
- verdverm 16d agoyou may reconsider OP's primary motivation based on another of their HN submissions https://news.ycombinator.com/item?id=49674396 https://news.ycombinator.com/item?id=49674396 my hunch is that an Ai has been validating their biases
- jwpapi 16d agoOkay I understand.. he’s up for the internet fame.
- _pdp_ 16d agoI don't really think there is much future for TypeSafe but I wish them well. In fact, I like them. Jev is just a reminder that you can use more "traditional" forms of AI (that are not LLMs) and still get remarkable results. We tend to forget that. That was a surprise and that is why it went viral.
- bjt12345 16d agoA question I have about Jev is, who are the "Service Providers" that they provide prompt information, and why is there no time limit on how long they store prompts? One of the Use Cases marketed is having Jev flag if personal information is contained in text. It's not a strong use case for it really.
- zamir_akimbekov 16d agobut now you got the attention. It is alright. Few remember Atanasof too.
- whywhywhywhy 16d ago"if you build it they will come" is a lie, marketing and branding matter and whats funny is often the biggest proponents of "if you build it" don't even realize the things they see going viral are marketed to them, they extremely naively think it got traction simply because it was good and they "built the right thing", not true.
- deleted 16d ago[deleted]
- slrainka 16d agoLook at the bright side. You can ride the Jev marketing, because at the end of the day, post prototype, data privacy is always going to be top of mind and people are already looking for Open Source alternatives because Jev proved the usecase in a simple way most people could understand.
- jerpint 16d agoAm I understanding correctly that the field has gone full circle and we are back to specialized classification models for domain specific tasks ?
- fernandezpablo 16d agoRunning against a labeled set (categorize support questions), jev gets 95% answers correct. Layla 48%
- i_remember_when 16d agoI did this a year before OP so how do you think I feel?
- deleted 16d ago[deleted]
- geocar 16d agoHi I want to explain that arxiv is not "publishing a paper" - it's a step up perhaps from putting it on your own website, but this is not what is meant by professional academics when they talk about "publishing" (even when they work for big AI companies). Your "papers" have only a single author and no current citations. It is not clear you are able to work with other people. The papers claim to offer results but no theory about why those results are the best possible. They read like sales whitepapers not scientific work. Reddit comments on the thread you linked to said they weren't able to reproduce your work. I'm not looking to buy magic beans for my sales team.
- zwaps 16d agoI am sympathetic but this buries the lede, hard. You are competitive with Jev only if you fine tune on the train dataset and calibrate per question. As much as I dislike literally everything about typesafes behavior, they have an API model that works on any problem without fine tuning, and that is the key. To be honest everyone who can finetune can likely finetune a BERT for a specific task and get similar results to yours. And that has been true for years The key to Jevs success is that it works without fine tuning
- lopuhin 16d agoIn the limitations you say that “Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.” - but that’s the whole point, you can obviously fine-tune specialized models but having a model follow you instructions and be promotable and fast makes it massively easier to use.
- baq 16d agoIn short you built a model, typesafe built a product
- tony_starkling 16d agoI doubt that this is as powerful as Jev. Of course, I don't think Jev is useful in the long run because it's virality stems from the fact that one of the builders worked at OpenAI pre chatGPT
- mochizou 16d agoI tried Jev a bit, and the zero-shot performance already felt good enough to be useful. If you find the right place for that tradeoff, I don’t really see why “you could fine-tune BERT” is much of a criticism.
- sporkland 15d ago> Seeing the hype online feels both validating and deeply frustrating. I worked on this literally one year back in March 2025. I spent months of hard work, sweat, and sleepless nights building it While I have deep empathy for what the author must be feeling. It feels tone deaf to me that as someone working in the model industry, likely using data sets scraped from other folks work, he doesn't see the hypocrisy in this. We've hit an interesting question in our society where a lot of us are getting tremendous value out of the uncredited and uncompensated use of people's hard work and sleepness nights. It seems naive of the author or anyone at this point to believe that the same won't happen to them. The question is whether we want this to be how our society operates? I can see a number of pros and cons and I know it's not a straight line from they trained on the works to derived works. But it's hard to imagine as humans how we will sustain broad motivation over a long period if we keep allowing this to happen.
- x3haloed 15d agoHere’s your medal.
- krenerd 15d agoboth RL part and Calibration honestly doesn't make sense to me. why does this use rl at all? i don't see why just CE cant resolve it? and why is it calibrated?
- vijay96238 15d agoNo of comments refer Jev on this thread 147! And with this, its 148!