Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
abdullin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
abdullin
1y ago
It takes deliberate practice to learn how to work with a new tool. I believe that AI+Coding is no different from this perspective. It usually takes senior engineers a few weeks just to start building an intuition of what is possible and wha
32.
▲
by
abdullin
1y ago
I heard an interesting story from an architect at a large software consultancy. They are using AI in their teams to manage legacy codebases in multiple languages. TLDR; it works for a codebase of 1M LoC. AI writes code a lot faster, comple
33.
▲
by
abdullin
1y ago
Compliance gaps / legal analysis is a pretty common theme in my community (meaning - it was mentioned 3-4 times by different teams). Here is how the approach usually looks like: 0. (the most painful step) Carefully parse all relevant d
34.
▲
by
abdullin
1y ago
Building something that people truly need - might not lead to huge sales right away, but I believe this to be a good long-term strategy. Sprint vs marathon. Just keep on pushing on it, and it will eventually work out.
35.
▲
by
abdullin
1y ago
When I was starting my community, I joined other similar communities and tried to be helpful there. No ads or links, just answering questions and supporting. People that were genuinely interested to learn more about the topic - opened my pr
36.
▲
by
abdullin
1y ago
There are way too many ads around "AI". Everybody else does that, frequently overwhelming people with too many promises of quick wins. I prefer to distinguish from this hype and reach people through other channels - good content,
37.
▲
by
abdullin
1y ago
Of course. Most of the AI cases (that turn out to be an actual success) focus around a few repeatable patterns and a limited use of "AI". Here are a few interesting ones: (1) Data extraction. E.g. extracting specs of electronic co
38.
▲
by
abdullin
1y ago
I have a course on building AI solutions in business (based on success stories from companies in Europe/USA). Sold ~400 seats so far, mostly through my community and word of mouth. No external ads or cold outreach. The process was clas
39.
▲
Benchmark of 110 RAG implementations on Annual Reports Q&A task
(abdullin.com)
1 points
by
abdullin
2y ago
|
1 comments
40.
▲
by
abdullin
2y ago
Results from the Enterprise RAG Challenge, comparing 110 experiments from 43 teams on building Retrieval-Augmented Generation (RAG) systems. The challenge required AI solutions to automatically answer 100 complex queries across 100 large an
41.
▲
by
abdullin
2y ago
Interesting read, thank you! Do you use any special tools to manage all these separate databases, track performance and debug problems?
42.
▲
Sama: "Insane thing: we are currently losing money on OpenAI pro subscriptions"
(twitter.com)
5 points
by
abdullin
2y ago
|
0 comments
43.
▲
by
abdullin
2y ago
This is an example of my prompt to o1 a few days ago. First request produced refactoring suggestions. Second request (yes, I like results, implement them) produced multiple files that have just worked. —- Take a look at this code from my mu
44.
▲
by
abdullin
2y ago
I would recommend giving a try to o1-preview in coding tasks like this one. It is one level above Claude 3.5 Sonnet, which currently is the most popular tool among my peers.
45.
▲
by
abdullin
2y ago
> For sure it will depend on use case, if you have fairly structured data or a clear domain-specific terminology to rely on Indeed. This works only in a subset of business domains for me: search and assistants within enterprise knowledge
46.
▲
by
abdullin
2y ago
It is impossible to fit all that information into the call. The whole point of RAG - we (somehow) retrieve only the relevant information and put it into the context to generate the answer.
47.
▲
by
abdullin
2y ago
It might depend on the case. My problem with similarity search - it is unpredictable. It can sometimes miss really obvious matches or pull completely irrelevant snippets. When this happens - this causes downstream hallucinations that are h
48.
▲
by
abdullin
2y ago
It really depends on the setup that the dev/ops at customer are more comfortable with. Elastic or PostgreSQL can be both fine. Personally for small cases (e.g. under 50k documents and 20GB of text) I like to use SQLite FTS, while linki
49.
▲
by
abdullin
2y ago
I’m consulting multiple teams on shipping LLM-driven business automation. So far I have seen only one case where fine-tuning a model really paid off (and didn’t just blow up the RLHF calibration and caused wild hallucinations). I would sugg
50.
▲
by
abdullin
2y ago
I’ve been building LLM-driven systems for customers for quite some time. We got tired of hallucinations from vector-based and hybrid RAGs last year, eventually arriving to the approach similar to yours. It is even called Knowledge Mapping [
51.
▲
by
abdullin
2y ago
Yep, exactly! My approach is similar - closed source benchmarks with prompts and tests from real LLM-driven products (mostly around boring business automation and enterprise workflows). Although it would be neat to upgrade the setup to work
52.
▲
by
abdullin
2y ago
Most of the blame that LLMs get is because they are used wrong. LLM is a good information transformation engine. But if I supply it with messy or wrong information in the context, then I’ll end up with hallucinations. This is why most of th
53.
▲
by
abdullin
2y ago
It all depends on the benchmark and the use case. If we are talking about the ability of models to follow instructions and carry out concrete tasks (as in products or inside RAG systems), then Gemini Pro 1.5 is currently on the eighth place
54.
▲
by
abdullin
2y ago
LocalLLaMA subreddit usually has some interesting benchmarks and reports. Here is one example, testing performance of different GPUs and Macs with various flavours of Llama: https://github.com/XiongjieDai/GPU-Benchmarks
55.
▲
by
abdullin
2y ago
There are too many variables at play, unfortunately. One can ran local LLMs even on RaspberryPi, although it will be horribly slow.
56.
▲
by
abdullin
2y ago
Yes, this can work. I’ve done that in a few cases. In fact, if you split data preprocessing in small enough steps, they could also be run on weaker LLMs. It would take a lot more time, but that is doable.
57.
▲
by
abdullin
2y ago
I think, that might come with the next GPT version. OpenAI seems to build in cycles. First they focus on capabilities, then they work on driving the price down (occasionally at some quality degradation)
58.
▲
by
abdullin
2y ago
I have a few LLM benchmarks that were extracted from real products. GPT-4o got slightly better overall. Ability to reason improved more than the rest.
59.
▲
by
abdullin
2y ago
IBM goes at great lengths to train models on clean data that has lower risk of copyright or legal issues attached. Just take a look at the model description. That data issue is important enough for some companies to pick mediocre model over
60.
▲
by
abdullin
2y ago
They are Ok-ish. Previous Granite models were on the level of first llama in my benchmarks. I’m expecting this version to be roughly comparable to llama 2
More ›