Agentic Hybrid RAG
Results reflect the tested configurations. Performance depends on models, retrieval settings, documents, and scoring.
81.33%
Overall accuracy
2
Refused
122
Correct
Evaluation analysis
Explore overall and per-category results for 150 questions. Test dates, model versions, retrieval settings, scoring details, and per-question outputs have not been disclosed.
Agentic Hybrid RAG vs Naive RAG vs closed-book LLM on the same FinanceBench set.
Results reflect the tested configurations. Performance depends on models, retrieval settings, documents, and scoring.
81.33%
Overall accuracy
2
Refused
122
Correct
Baseline
Naive RAG
The basic retrieval-augmented configuration in this evaluation.
Overall accuracy 21.33%
Baseline
Closed-book LLM (no docs)
No retrieved documents are supplied when answering.
Overall accuracy 24%
Naive RAG
21.33%Closed-book LLM (no docs)
24%Naive RAG
69.33%Closed-book LLM (no docs)
42%| System | Correct | Incorrect | Refused | Overall accuracy |
|---|---|---|---|---|
| Agentic Hybrid RAGOurs | 122 | 26 | 2 | 81.33% |
| Naive RAG | 32 | 14 | 104 | 21.33% |
| Closed-book LLM (no docs) | 36 | 51 | 63 | 24% |
The eval set has 50 questions each in domain, metrics, and novel types. Below compares overall accuracy across all three systems.
Domain
Metrics
Novel
All questions
Correct answers, incorrect answers, refusals, question counts, and both accuracy measures match the summary CSV.
| System | Question type | Correct | Incorrect | Refused | Question count | Overall accuracy | Accuracy among attempted questions |
|---|---|---|---|---|---|---|---|
| Agentic Hybrid RAGOurs | All questions | 122 | 26 | 2 | 150 | 81.33% | 82.43% |
| Agentic Hybrid RAG | Domain | 41 | 8 | 1 | 50 | 82% | 83.67% |
| Agentic Hybrid RAG | Metrics | 44 | 6 | 0 | 50 | 88% | 88% |
| Agentic Hybrid RAG | Novel | 37 | 12 | 1 | 50 | 74% | 75.51% |
| Naive RAG | All questions | 32 | 14 | 104 | 150 | 21.33% | 69.57% |
| Naive RAG | Domain | 12 | 7 | 31 | 50 | 24% | 63.16% |
| Naive RAG | Metrics | 1 | 1 | 48 | 50 | 2% | 50% |
| Naive RAG | Novel | 19 | 6 | 25 | 50 | 38% | 76% |
| Closed-book LLM (no docs) | All questions | 36 | 51 | 63 | 150 | 24% | 41.38% |
| Closed-book LLM (no docs) | Domain | 19 | 15 | 16 | 50 | 38% | 55.88% |
| Closed-book LLM (no docs) | Metrics | 10 | 18 | 22 | 50 | 20% | 35.71% |
| Closed-book LLM (no docs) | Novel | 7 | 18 | 25 | 50 | 14% | 28% |
Review correct answers, incorrect answers, and refusals for each question type.
All questions
Full eval set across domain, metrics, and novel questions.
81.33%
Accuracy among attempted questions 82.43% · Question count 150
Domain
Questions grounded in the domain knowledge base.
82%
Accuracy among attempted questions 83.67% · Question count 50
Metrics
Questions that need numbers, rates, or computed facts.
88%
Accuracy among attempted questions 88% · Question count 50
Novel
Questions outside the usual pattern—harder to fake with memory.
74%
Accuracy among attempted questions 75.51% · Question count 50
Overall accuracy = correct ÷ all questions. Accuracy among attempted questions = correct ÷ (correct + incorrect). Refusals count as not correct. Source: accuracy_summary.csv.
Firmground AI
Choose a task your team already knows, and see how Firmground AI handles it and what it delivers.
Existing users can sign in. Book a demo to discuss enterprise integrations and deployment.