Papers
arxiv:2607.27853

FinanceHarness: Autonomous Financial Deep Research Framework

Published on Aug 7
· Submitted by
Han
on Aug 6
Authors:
,
,
,
,
,
,
,
,

Abstract

FinanceHarness automates specialized financial deep research through structured agent workflows and verifiable benchmarks, while FinanceGym provides rigorous evaluation criteria revealing substantial room for improvement.

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.

Community

Paper submitter

We built a point-in-time financial deep research benchmark, featuring questions and rubrics generated through a rigorous, quality-controlled data pipeline. Additionally, we contracted financial experts to validate our data, with each spending an average of 1.2 hours on this meticulous review process. Leading LLMs such Opus-5 with our Finance Harness only score 44.9% on our leaderboard, showcasing the significant challenge our benchmark presents. We invite everyone to contribute to our leaderboard!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.27853 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.27853 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.