Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
airylizard
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
A design brief as one image: can a 27B vision model build the animation?
(automationoptimization.github.io)
3 points
by
airylizard
11d ago
|
0 comments
2.
▲
Show HN: Patient Glue a more affordable SMS solution for healthcare that I built
(patientglue.com)
1 points
by
airylizard
1y ago
|
0 comments
3.
▲
Think Before You Speak – Exploratory Forced Hallucination Study [pdf]
(github.com)
2 points
by
airylizard
1y ago
|
0 comments
4.
▲
Evaluating AI Agents with Azure AI Evaluation
(techcommunity.microsoft.com)
5 points
by
airylizard
1y ago
|
0 comments
5.
▲
by
airylizard
1y ago
Exactly what leads to inaccurate output in LLM's. The semantic interpretation of each individual token isn't the same between us and it. "Interpretation", we likely define accuracy and reliability pretty closely the same
6.
▲
by
airylizard
1y ago
Are you continuing research? Is there somewhere we can follow along?
7.
▲
by
airylizard
1y ago
The fact that embeddings from different models can be translated into a shared latent space (and back) supports the notion that semantic anchors or guides are not just model-specific hacks, but potentially universal tools. Fantastic read, t
8.
▲
by
airylizard
1y ago
As more and more people use brute force loops to make their AI agents more reliable, this hidden inference giant will only continue to grow. This is why I put my framework together, using just 2 passes as opposed to n+ can increase accuracy
9.
▲
by
airylizard
1y ago
love it. any llm can be made to perform reliably and accurately which is the biggest pre-requisite when it comes to creating an "AI Agent". I think this gives people the opportunity to start somewhere because they can leverage mul
10.
▲
by
airylizard
1y ago
The data "supply chain" has already surged ahead of production elsewhere. Companies aren't just passively taking what's out there, they actively harvest highly curated content, benefiting even further when we voluntarily
11.
▲
by
airylizard
1y ago
I like the idea, TSCE framework should make the individual agents more reliable and deterministic: https://github.com/AutomationOptimization/tsce_demo
12.
▲
by
airylizard
1y ago
Right, the 4.1 training checkpoint hasn’t moved. What has moved is the glue on top: decoder heuristics / safety filters / logit-bias rules that OpenAI can hot-swap without re-training the model. Those “serving-layer” tweaks are wh
13.
▲
by
airylizard
1y ago
Hey, thanks for kicking the tires! The run you’re describing was done in mid-April, right after GPT-4.1 went live. Since then OpenAI has refreshed the weights behind the “gpt-4.1” alias a couple of times, and one of those updates fixed the
14.
▲
by
airylizard
1y ago
The test isn't for how well an LLM can find or replace a string. It's for how well it can carry out given instructions... Is that not obvious?
15.
▲
by
airylizard
1y ago
Why I came up with TSCE(Two-Step Contextual Enrichment). +30pp uplift when using GPT-35-turbo on a mix of 300 tasks. Free open framework, check the repo try it yourself https://github.com/AutomationOptimization/tsce_dem
16.
▲
by
airylizard
1y ago
1. What TSCE is in one breath Two deterministic forward-passes. 1. The model is asked to emit a hyperdimensional anchor (HDA) under high temperature. 2. The same model is then asked to answer while that anchor is prepended to the original p
17.
▲
TSCE and HyperDimensional Anchors: Making AI agents/workflows reliable at scale
(github.com)
3 points
by
airylizard
1y ago
|
1 comments
18.
▲
Show HN: TSCE – Think Before You Speak (Two-Step Contextual Enrichment for LLMs)
(github.com)
3 points
by
airylizard
1y ago
|
0 comments