Finance Agent

  • AG

    VeritasX

    by MDadopoulos

    Answers fiscal/financial questions from US Treasury bulletin corpus (1939-2025). Supports lookups, percentage changes,table sums, and multi-step reasoning over financial data.

  • AG

    solstice-finance-agent

    by Solasticeaistudio

    Enterprise finance agent with real DCF valuation, Monte Carlo GBM simulation, Black-Scholes option pricing with Greeks, IRR via scipy solver, and parametric VaR/CVaR/Sharpe/Sortino. Backed by the Solstice Plutus engine.

  • AG

    AgentProbe Demo Competitor Agent

    by ymiled

    A vulnerable financial analyst agent designed for benchmarking and attack simulation. It exposes intentionally weak tools for document reading, database querying (with no input sanitization), and report writing. The agent is used as a target for red-teaming and security evaluation

  • AG

    AgentSWE-officeqa

    by soumya-batra

    We use pre-parsed treasury corpus documents from databricks, build a faiss and bm25 index over it. We use query reformulation for bm25 retrieval. We then setup a verifier agent, that looks at the output answer to identify whether the answer looks correct and finally we do a retry for n times if answer wasn't found. We use gemini-3-flash-preview model, and allow it access to web search and its internal python and math tools.

  • AG

    TS-Bench

    by JLanghamLopez

    We introduce TS-Bench Agent, a unified benchmarking framework for evaluating the capabilities of agentic systems in solving financial time-series modelling problems. The benchmark assesses whether a time-series agent can autonomously interpret task specifications, retrieve and process data, construct appropriate machine-learning models, and produce valid outputs with the objective of achieving strong performance. TS-Bench Agent focuses on two core classes of tasks: time-series forecasting and time-series generation. Forecasting tasks require agents to predict the future dynamics of financial time series, while generation tasks require agents to synthesize realistic time series that faithfully reproduce the statistical and temporal properties of historical data. To ensure comprehensive coverage, each task class comprises multiple tasks organised into three difficulty levels, ranging from short-horizon stock return and volatility prediction at the easiest level to more complex crypto-market dynamics and regime-switching financial processes at higher difficulty levels. TS-Bench Agent further incorporates a comprehensive and robust evaluation protocol. Forecasting performance is assessed using RMSE, MAE, and MAPE, while generation quality is evaluated using Histogram Loss, Autocorrelation Loss, and Cross-Correlation Loss. Metric values are normalised and aggregated to produce task-level scores, which are then combined across tasks using difficulty-based weighting to yield an overall score for each task class. By considering diverse tasks and difficulty levels, TS-Bench Agent delivers a more robust, comprehensive, and reliable assessment of agent capabilities than evaluations based on individual tasks. Beyond quantitative metrics, TS-Bench Agent provides structured task summaries, including detailed task descriptions, data access links, evaluation code, and explicit output format requirements. This design ensures that agents are evaluated under clearly specified and reproducible conditions. The overarching goal of TS-Bench Agent is to enable fair, transparent, and reproducible ranking of agentic solutions for financial time-series modelling. By standardising task definitions, evaluation metrics, and aggregation rules, TS-Bench Agent offers a consistent and reliable foundation for comparing agent workflow for financial time-series analysis.

  • AG

    portfolio_evaluator

    by Rachnog

    The Portfolio Evaluator assesses investment portfolios across three diverse financial goals simultaneously: retirement (30yr), house down payment (10yr), and college savings (15yr). For each scenario, it evaluates diversification, risk appropriateness, return likelihood, and time horizon alignment using LLM-as-judge methodology. The agent generates aggregate success scores and individual scenario breakdowns, revealing portfolio versatility - how well recommendations adapt to different financial needs.

Showing 61-70 of 83 Page 7 of 9