3 ms·
I am confused by the benchmarks you provided. (1) The "baseline" comparison was to PyPDF + Naive RAG. For the LlamaParse evaluation, you appear to have used a
by brendanashworth 3y ago
I am confused by the benchmarks you provided.
(1) The "baseline" comparison was to PyPDF + Naive RAG. For the LlamaParse evaluation, you appear to have used a different RAG pipeline, called "recursive retrieval." Why not use the same pipeline to demonstrate the improvement from LlamaParse? Can you share the code to your evaluation for LlamaParse?
(2) I ran the benchmark for the PyPDF + Naive RAG solution, directly copying the code on the linked LlamaIndex repo [1]
I got very different numbers:
mean_correctness_score 3.941
mean_relevancy_score 0.826
mean_faithfulness_score 0.980
You reported:
mean_correctness_score 3.874
mean_relevancy_score 0.844
mean_faithfulness_score 0.667
Notably, the faithfulness score I measured for the baseline solution was actually higher than that reported for your proprietary LlamaParse based solution.
[1] https://github.com/run-llama/llama-hub/tree/main/llama_hub/llama_datasets/10k/uber_2021 https://github.com/run-llama/llama-hub/tree/main/llama_hub/l...
- freezed8 3y ago(jerry here) Thanks for running through the benchmark! Just to clarify some things: (1) The idea is that LlamaParse's markdown representation lends itself to the rest of LlamaIndex advanced indexing/retrieval abstractions. Recursive retrieval is a fancy retrieval method designed to model documents with embedded objects, but depends on good PDF parsing. Naive PyPDF parsing can't be used with recursive retrieval. Our goal is to demonstrate the e2e RAG capabilities of LlamaParse + advanced retrieval vs. what you can build with a naive PDF parser. (2). Since we use LLM-based evals, your correctness and relevancy metric look to be consistent and within margin of error (and lower than our llamaparse metrics). The faithfulness score seems way off though and quite high from your side, so not sure what's going on there. maybe hop in our discord and share the results in our channel?