Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mingtianzhang
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
Show HN: We Built a Chat of Stanford's CS229 Course Notes
4 points
by
mingtianzhang
3mo ago
|
0 comments
2.
▲
by
mingtianzhang
6mo ago
Hi, thanks for the feedback. Our biggest differentiation is that OpenKB can handle long PDFs and images—something that isn’t trivial or easily “vibe-coded.” We will try to find some ways to compare with other projects! Thanks for the valuab
3.
▲
Show HN: Open KB: Open LLM Knowledge Base
6 points
by
mingtianzhang
6mo ago
|
2 comments
4.
▲
ClawdReview – OpenReview for AI Agents
5 points
by
mingtianzhang
8mo ago
|
0 comments
5.
▲
Show HN: ClawdReview – OpenReview for AI Agents
3 points
by
mingtianzhang
8mo ago
|
0 comments
6.
▲
by
mingtianzhang
11mo ago
VLM can already process both the document images and the query to produce an answer directly. Do we still need the intermediate OCR step?
7.
▲
Do we still need OCR? An implementation of a pure vision-based agent
(pageindex.ai)
7 points
by
mingtianzhang
11mo ago
|
1 comments
8.
▲
by
mingtianzhang
11mo ago
We discuss the limitations of the classic OCR pipeline and provide a pure vision-based RAG system for document analysis ( https://github.com/VectifyAI/PageIndex/blob/main/cookbook/vi... ) Any feedback
9.
▲
by
mingtianzhang
11mo ago
We actually don't need OCR: https://pageindex.ai/blog/do-we-need-ocr
10.
▲
Do We Still Need OCR?
(pageindex.ai)
4 points
by
mingtianzhang
11mo ago
|
2 comments
11.
▲
by
mingtianzhang
11mo ago
This blog examines the inherent limitations of the current OCR pipeline in the context of document question-answering systems from an information-theoretic perspective and discusses why a direct, vision-based approach can be more effective.
12.
▲
Reasoning-based RAG for long document question answering
(pageindex.ai)
1 points
by
mingtianzhang
1y ago
|
1 comments
13.
▲
by
mingtianzhang
1y ago
PageIndex Chat is the world's first human-like long-document AI analyst. You can upload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser. Unlike traditional RAG or &qu
14.
▲
PageIndex Chat – Human-Like Long Document AI Analyst
(pageindex.ai)
6 points
by
mingtianzhang
1y ago
|
1 comments
15.
▲
by
mingtianzhang
1y ago
PageIndex Chat is the world's first human-like long-document AI analyst. You can upload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser. Unlike traditional RAG or &qu
16.
▲
PageIndex Chat – Human-Like Long Document AI Analyst
(pageindex.ai)
4 points
by
mingtianzhang
1y ago
|
1 comments
17.
▲
by
mingtianzhang
1y ago
PageIndex Chat is the world's first human-like long-document AI analyst. You can pload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser. Unlike traditional RAG or &quo
18.
▲
Show HN: In-Context Index for In-Context Retrieval
(github.com)
5 points
by
mingtianzhang
1y ago
|
0 comments
19.
▲
DeepMind's paper reveals Google's new direction on RAG: In-Context Retreival
(arxiv.org)
6 points
by
mingtianzhang
1y ago
|
1 comments
20.
▲
by
mingtianzhang
1y ago
Instead of relying on vector databases, DeepMind proposes: 1. The LLM itself selects the most relevant documents — no vector database needed. 2. The selected documents are then placed directly into the context for generation. This kind of i
21.
▲
From Claude Code to Agentic RAG
(vectifyai.notion.site)
3 points
by
mingtianzhang
1y ago
|
0 comments
22.
▲
From Claude Code to PageIndex: The Rise of Agentic Retrieval
(vectifyai.notion.site)
3 points
by
mingtianzhang
1y ago
|
0 comments
23.
▲
From Claude Code to PageIndex: The Rise of Agentic Retrieval
(vectifyai.notion.site)
8 points
by
mingtianzhang
1y ago
|
0 comments
24.
▲
by
mingtianzhang
1y ago
Hi, thanks for your inspiring questions. 1. What happens when the TOC is too long? -- This is why we choose the tree structure. If the ToC is too long, it will do a hierarchy search, which means search over the father level nodes first and
25.
▲
Show HN: PageIndex for Reasoning-Based RAG
(vectifyai.notion.site)
6 points
by
mingtianzhang
1y ago
|
0 comments
26.
▲
by
mingtianzhang
1y ago
The current OCR approach typically relies on a Vision-Language Model (VLM) to convert a table into a JSON structure. However, a table inherently has a 2D spatial structure, while Large Language Models (LLMs) are optimized for processing 1D
27.
▲
Show HN: A Vectorless LLM-Native Document Index Method
(github.com)
14 points
by
mingtianzhang
1y ago
|
3 comments
28.
▲
by
mingtianzhang
1y ago
Thanks, any feedback is welcome!
29.
▲
by
mingtianzhang
1y ago
Thanks for the reminder, I have edited the comment.
30.
▲
by
mingtianzhang
1y ago
Edited version: We try to solve a similar problem to put long documents in context. We built an MCP for Claude to allow you to put long PDFs in your context window that go beyond the context limits: https://pageindex.ai/mcp
More ›