Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
renchuw
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Show HN: InverSQL: Build SQL interactively in inverse
(github.com)
6 points
by
renchuw
27d ago
|
0 comments
2.
▲
Show HN: Creating SQL queries with decision trees
(inversql.rentruewang.com)
3 points
by
renchuw
4mo ago
|
1 comments
3.
▲
by
renchuw
4mo ago
Create SQL by over fitting decision tree on data, then optimize the boolean representation.
4.
▲
by
renchuw
3y ago
Well, this method is based on the assumption that embeddings can accurately represent the texts and their structural relations are preserved. So long as you have all the random seeds fixed, I think reproduction should be straight forward.
5.
▲
by
renchuw
3y ago
Thanks for the feedback! The reason the "code" part is more complete than the "research" part is because I originally planned for it to just be a hobby project and only very later on decided to perhaps try to be serious
6.
▲
by
renchuw
3y ago
Correct.
7.
▲
by
renchuw
3y ago
Hi, OP here. I would kind of have to disagree here. You raised some interesting points, but I don't think something can be qualified as *moat* if it is overcome-able by just sharing the use cases. For example, we all know Google's
8.
▲
by
renchuw
3y ago
This would be an inner loop process. However, the selection is way faster than LLMs so it shouldn't be noticable (hopefully).
9.
▲
by
renchuw
3y ago
Hi, OP here. I would say not really because the goals are different. Although both uses retrieval techniques, RAG wants to augment your query with factual information, where here we retrieve in order to evaluate on as few queries as possibl
10.
▲
by
renchuw
3y ago
I designed 2 modes in the project, exploration mode and exploitation mode. Exploration mode uses entropy search to explore the latent space (used for evaluating the LLM on the selected corpus to evaluate), and eploitation mode is used t
11.
▲
by
renchuw
3y ago
Perhaps I should clarify it in the project README. It's the phase to evaluate how well your model is performing. So the pipeline goes training -> evaluation -> deployment (inference) corresponding to the datasets in supervised tr
12.
▲
by
renchuw
3y ago
Fair question. Evaluate refers to the phase after training to check if the training is good. Usually the flow goes training -> evaluation -> deployment (what you called inference). This project is aimed for evaluation. Evaluation can
13.
▲
by
renchuw
3y ago
Hi, OP here. It's not 10 times faster inference, but faster evaluation. You use evaluation on a dataset to check if your model is performing well. This takes a lot of time (might be more than training if you are just finetuning a pre-t
14.
▲
by
renchuw
3y ago
Hi, OP here. So you evaluate LLMs on corpuses to evaluate their performance right? Bayesian optimization is here to select points (in the latent space) and tell the LLM where to evaluate next. To be precise, entropy search is used here (cou
15.
▲
by
renchuw
3y ago
Hi, OP here, sorry for late reply. I am not actually "evaluating", but rather using the "side effects" of bayesian optimization that allows zoning in/out on some regions on the latent space. Since embedders are so f
16.
▲
by
renchuw
3y ago
Side note: OP here, I came up with this cool idea because I was chatting with a friend about how to make LLM evaluations fast (which is so painfully slow on large datasets) and realized that somehow no one has tried it. So I decided to give
17.
▲
Show HN: Faster LLM evaluation with Bayesian optimization
(github.com)
131 points
by
renchuw
3y ago
|
43 comments