4 ms·
I don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is t
by edot 15d ago
I don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is the number odd?”
Answer: 9% chance, with 91% confidence.
Heh???
Ok, even worse. 75% chance a coin landed heads up?
State: I flipped a coin.
Question:
{
"noul_result": {
"type": "noul",
"instructions": "Did the coin land heads up?"
},
"choice_result": {
"type": "choice",
"instructions": "Determine if the coin landed heads or tails up.",
"criteria": {
"heads": "the coin landed heads up",
"tails": "the coin landed tails up"
}
}
}
Ran on: https://huggingface.co/spaces/convaiinnovations/laya-demo https://huggingface.co/spaces/convaiinnovations/laya-demo
Result:
{
"model": "laya",
"answers": {
"noul_result": {
"type": "noul",
"noul": 0.6839,
"rl_agent": {
"act_probability": 1.0
}
},
"choice_result": {
"type": "choice",
"choice": "heads",
"probabilities": {
"heads": 0.7407,
"tails": 0.2593
},
"confidence": 0.1743,
"rl_agent": {
"act_probability": 1.0
}
}
},
"usage": {
"input_tokens": 76,
"output_tokens": 0
},
"latency_ms": 93.8
}
Trying to be even more good-faith:
State: "A fair coin was flipped once. The result was not observed.
No other information about the outcome is available."
Questions:
{
"noul_result": {
"type": "noul",
"instructions": "Given only the supplied state, what is the probability that the coin landed heads up?"
},
"choice_result": {
"type": "choice",
"instructions": "Given only the supplied state, determine which outcome occurred.",
"criteria": {
"heads": "the coin landed heads up",
"tails": "the coin landed tails up"
}
}
}
Result:
{
"model": "laya",
"answers": {
"noul_result": {
"type": "noul",
"noul": 0.1265,
"rl_agent": {
"act_probability": 1.0
}
},
"choice_result": {
"type": "choice",
"choice": "tails",
"probabilities": {
"heads": 0.2522,
"tails": 0.7478
},
"confidence": 0.1853,
"rl_agent": {
"act_probability": 1.0
}
}
},
"usage": {
"input_tokens": 123,
"output_tokens": 0
},
"latency_ms": 154.5
}
- bensyverson 15d agoThis is not a good faith test of the system.
- deleted 15d ago[deleted]
- edot 15d agoBut it's hallucination-free, isn't it?
- usagisushi 15d agoyeah, technically. (/s) python3 - <<'EOF' import json, urllib.request body = json.dumps({ "state": "The car wash is only 100 meters away from my house.", "model": "jev-1.13-free", "questions": {"q": {"type": "choice", "instructions": "Should I drive or walk to the car wash?", "criteria": {"drive a car": None, "walk": None}}} }).encode() req = urllib.request.Request("https://opencode.ai/zen/v1/systemone", data=body, headers={"Content-Type": "application/json", "User-Agent": "opencode/1.18.31"}) with urllib.request.urlopen(req, timeout=60) as r: print(json.dumps(json.load(r)["answers"]["q"], indent=2)) EOF { "type": "choice", "choice": "walk", "confidence": 0.66, "probabilities": { "walk": 0.83, "drive a car": 0.17 } }
- bensyverson 15d agoI guess we’ve just reached the point where everyone has to state the obvious, and common sense is extremely uncommon. So here goes: you should not use an AI model to validate a claim which is trivial to calculate deterministically. That is (obviously?) not what a model like Jev is for, thus it is not a good test of Jev.
- hmokiguess 15d agoMaybe it’s the fact that “the number” could refer to both 6 and 3 to this model?
- prometheus1992 15d agotry this model on HF - https://huggingface.co/MoritzLaurer/deberta-v3-large-zeroshot-v2.0?candidate_labels=the+number+is+odd%2C+the+number+is+even&multi_class=false&text=a+6+sided+die+rolled+a+3 https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho... a 6 sided die rolled a 3 possible class names - the number is odd, the number is even result: the number is odd 0.945 the number is even 0.055 as someone else said, that 0.055 is probably bc of 6 and 3 being there.
- ksymph 15d agoI don't think calculating mathematical odds from natural language is the sort of problem this is trying to solve. A typical LLM hooked up to a calculator would be more appropriate for that. Jev (and similar) is more for data processing and sentiment analysis. Moderation, search engines, that sort of thing. Jev has a page of proposed use cases where you can get an idea of what they're going for: https://docs.typesafe.ai/concepts/use-case-map https://docs.typesafe.ai/concepts/use-case-map
- jwpapi 15d agoJev says you should restate state in the question and I tried it: { "decision": { "type": "noul", "instructions": "Is the rolled number in state odd?" }, "question": { "type": "noul", "instructions": "Is the number odd?" }, "question-3": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the number odd?" }, "question-4": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the rolled number odd?" } } => decision,0.168,0.83 question,0.141,0.86 question-3,0.029,0.97 question-4,0.021,0.98 so im confused too.. A weakness with numbers?
- bensyverson 15d agoBreaking news: small language models struggle with math
- hbrn 15d agoBut didn’t you hear? > Jev is neither small nor an LLM
- bensyverson 14d agoIt's either a small language model or a large language model (LLM). It's not a generative model, but neither is BERT, which is also a language model.