Evaluation analysis

Financial Q&A evaluation results

Explore overall and per-category results for 150 questions. Test dates, model versions, retrieval settings, scoring details, and per-question outputs have not been disclosed.

Overall results

Agentic Hybrid RAG vs Naive RAG vs closed-book LLM on the same FinanceBench set.

Ours

Agentic Hybrid RAG

Results reflect the tested configurations. Performance depends on models, retrieval settings, documents, and scoring.

81.33%

Overall accuracy

2

Refused

122

Correct

Baseline

Naive RAG

The basic retrieval-augmented configuration in this evaluation.

Overall accuracy 21.33%

Baseline

Closed-book LLM (no docs)

No retrieved documents are supplied when answering.

Overall accuracy 24%

  • Agentic Hybrid RAG recorded 81.33% accuracy.
  • Naive RAG recorded 104 refusals out of 150 questions in this evaluation.
  • The closed-book model answers without retrieved documents and recorded 24.00% accuracy.

Overall accuracy

Agentic Hybrid RAGOurs

81.33%

Naive RAG

21.33%

Closed-book LLM (no docs)

24%

Overall refusal rate

Agentic Hybrid RAG2 refusals

1.33%

Naive RAG

69.33%

Closed-book LLM (no docs)

42%
SystemCorrectIncorrectRefusedOverall accuracy
Agentic Hybrid RAGOurs12226281.33%
Naive RAG321410421.33%
Closed-book LLM (no docs)36516324%

Breakdown by question type

The eval set has 50 questions each in domain, metrics, and novel types. Below compares overall accuracy across all three systems.

Our solution · overall accuracy by type

Domain

82%

Metrics

88%

Novel

74%

All questionsOurs

81.33%

Overall accuracy by question type

  • Hybrid
  • Naive
  • Closed

Domain

Hybrid
82%
Naive
24%
Closed
38%

Metrics

Hybrid
88%
Naive
2%
Closed
20%

Novel

Hybrid
74%
Naive
38%
Closed
14%

All questions

Hybrid
81.33%
Naive
21.33%
Closed
24%

Results by system and question type

Correct answers, incorrect answers, refusals, question counts, and both accuracy measures match the summary CSV.

SystemQuestion typeCorrectIncorrectRefusedQuestion countOverall accuracyAccuracy among attempted questions
Agentic Hybrid RAGOursAll questions12226215081.33%82.43%
Agentic Hybrid RAGDomain41815082%83.67%
Agentic Hybrid RAGMetrics44605088%88%
Agentic Hybrid RAGNovel371215074%75.51%
Naive RAGAll questions321410415021.33%69.57%
Naive RAGDomain127315024%63.16%
Naive RAGMetrics1148502%50%
Naive RAGNovel196255038%76%
Closed-book LLM (no docs)All questions36516315024%41.38%
Closed-book LLM (no docs)Domain1915165038%55.88%
Closed-book LLM (no docs)Metrics1018225020%35.71%
Closed-book LLM (no docs)Novel718255014%28%

Our solution · summary by type

Review correct answers, incorrect answers, and refusals for each question type.

All questions

Full eval set across domain, metrics, and novel questions.

81.33%

Accuracy among attempted questions 82.43% · Question count 150

Domain

Questions grounded in the domain knowledge base.

82%

Accuracy among attempted questions 83.67% · Question count 50

Metrics

Questions that need numbers, rates, or computed facts.

88%

Accuracy among attempted questions 88% · Question count 50

Novel

Questions outside the usual pattern—harder to fake with memory.

74%

Accuracy among attempted questions 75.51% · Question count 50

Overall accuracy = correct ÷ all questions. Accuracy among attempted questions = correct ÷ (correct + incorrect). Refusals count as not correct. Source: accuracy_summary.csv.

Start with one business task using an AI Agent

Firmground AI

Choose a task your team already knows, and see how Firmground AI handles it and what it delivers.

  • Process business documents and data
  • Connect business tools and carry out the task
  • Reuse workflows, with approval before key actions

Existing users can sign in. Book a demo to discuss enterprise integrations and deployment.