Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
codelion
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
Pivotal Token Search (PTS): Targeting Critical Decision Points in LLM Training
(huggingface.co)
2 points
by
codelion
1y ago
|
1 comments
32.
▲
by
codelion
1y ago
It is by design. OpenAI is not going to reveal any architectural innovation they have made in their own commercial models.
33.
▲
Internal Coherence Maximization(ICM): Label-Free Unsupervised Training Framework
(github.com)
1 points
by
codelion
1y ago
|
0 comments
34.
▲
Unsupervised Model Improvement via Internal Coherence Maximization
(huggingface.co)
1 points
by
codelion
1y ago
|
0 comments
35.
▲
Show HN: PTS Library – Analyze LLM reasoning through "thought anchors"
2 points
by
codelion
1y ago
|
0 comments
36.
▲
Automated Discovery of High-Performance GPU Kernels with OpenEvolve
(huggingface.co)
4 points
by
codelion
1y ago
|
0 comments
37.
▲
Adaptive Classifier: Dynamic Text Classification with Continuous Learning
(huggingface.co)
1 points
by
codelion
1y ago
|
0 comments
38.
▲
Show HN: DeepThink Plugin – Bring Gemini 2.5's parallel reasoning to open models
4 points
by
codelion
1y ago
|
0 comments
39.
▲
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques
(arxiv.org)
1 points
by
codelion
1y ago
|
0 comments
40.
▲
by
codelion
1y ago
You can run in two modes, by default you run in the inference mode without learning. So, the changes you made will be used. If you switch to learning mode then the strategies are updated/refined and merged based on a config that you ca
41.
▲
by
codelion
1y ago
We do not allow the strategies to keep growing there is a refinement phase where we refine and merge existing strategies. The experiments were run with this config - https://github.com/codelion/optillm/blob/ma
42.
▲
by
codelion
1y ago
Re-reading the problem apparently works well - https://arxiv.org/abs/2309.06275 Here the system seems to have discovered this strategy by itself. The prompts are generic because during learning there is a part to refin
43.
▲
by
codelion
1y ago
We have some examples in the plugin README: https://github.com/codelion/optillm/tree/main/optillm/plugin... E.g. This was the strategy discovered by optiLLM for solving word problems: *Refined Stra
44.
▲
by
codelion
1y ago
Optillm works with llama.cpp but this approach is implemented as a decoding strategy in PyTorch so at the moment you will need to use the local inference server in optillm to use it.
45.
▲
by
codelion
1y ago
Thanks for checking this out! A few additional details that didn't fit in the main post: The system maintains two separate limits: a storage limit (max 10 strategies per problem type in the database) and an inference limit (max 3 strat
46.
▲
Show HN: System Prompt Learning – LLMs Learn Problem-Solving from Experience
48 points
by
codelion
1y ago
|
13 comments
47.
▲
by
codelion
1y ago
If we try to benchmark GPQA-Diamond with DeepSeek-R1 in the suggested configuration of 0.6 temp and 32k max_tokens and say if every instance takes the maximum tokens it will require 6.4 M tokens. Which without batching on a single H100 at 8
48.
▲
by
codelion
1y ago
This is an interesting idea, I hadn't thought of it. It is worth experimenting I am not aware of anyone else trying it yet.
49.
▲
by
codelion
1y ago
Hey, yes the reported results do not restrict any time limit or token limit for the benchmarks. We run our baseline with the same config 0.6 temp and max_token 32k but we set a timeout after 600 secs. Otherwise it would take forever to benc
50.
▲
by
codelion
1y ago
We use an adaptive classifier to learn how many tokens the model takes to respond correctly on a known dataset. I used the https://huggingface.co/adaptive-classifier/llm-router for experiments it is based on distilbert
51.
▲
by
codelion
1y ago
The short answer is in general yes it helps improve the accuracy, there is a whole line of work on self consistency and critique that supports it. Many of those approaches are already implemented in optillm.
52.
▲
by
codelion
1y ago
Yes, the goal here is to avoid overthinking and be as efficient as possible in terms of the minimal tokens required to solve a query. Often, queries that require too many tokens are unlikely to lead to correct answers anyways otherwise they
53.
▲
by
codelion
1y ago
Query complexity in this context is based on how many tokens it took for the model to respond to a query correctly based on a ground truth dataset like GSM8k. The adaptive classifier learns over this dataset and then we use it at inference
54.
▲
by
codelion
1y ago
This sounds like an interesting idea, can you elaborate more may be with a concrete example. I am wondering if this can be implemented easily as a plugin in optillm.
55.
▲
by
codelion
1y ago
Yes, we started with the idea of trying to replicate similar control on thinking processes for open reasoning models. They also announced the Deep Think approach at IO which goes even further and combines parallel CoTs at inference.
56.
▲
by
codelion
1y ago
The motivation for AutoThink came from watching how current reasoning models waste computation - they spend the same amount of "thinking time" on "what's 2+2?" as they do on complex mathematical proofs. This seemed
57.
▲
Show HN: AutoThink – Boosts local LLM performance with adaptive reasoning
397 points
by
codelion
1y ago
|
68 comments
58.
▲
by
codelion
1y ago
Checked my openrouter stats, it took ~3M tokens but that involved quite a few runs of various experiments.
59.
▲
by
codelion
1y ago
You can try an open-source implementation - https://github.com/codelion/openevolve
60.
▲
by
codelion
1y ago
I actually managed to replicate the new SOTA for circle packing in unit squares as found in the alphaevole paper - 2.635 for 26 circles in a unit square. Took about 800 iterations to find the best program which itself uses an optimisation p
More ›