BuySideBench 1.0 · Itoflow Research
A benchmark for quantitative portfolio decisions
BuySideBench gives an AI system an investor request, point-in-time market data, trading costs, and portfolio constraints. The score is based on the submitted portfolio action.
What the benchmark tests
Research under uncertainty and investor constraints
The tasks require systems to interpret evidence, account for costs and constraints, and make a portfolio decision without knowing which market outcome will occur.
Itoflow researches investment questions and produces portfolio decisions subject to investor constraints. BuySideBench tests that work in two settings: complete initial instructions and assignments where the system must ask for missing investor constraints. Itoflow is evaluated on the same tasks and by the same scorer as the other systems shown below.
A task in practice
The system had enough evidence, but not enough investor information.
One assignment asks the system to improve an eight-equity portfolio after new earnings and guidance. It has the data needed to research the companies, but the brief does not reveal what portfolio the investor will accept. That distinction is the center of the benchmark.
What the system could research
Signals
Point-in-time corporate-event allocation
Improve an eight-equity sleeve after the latest earnings and guidance updates using only information available at the decision time.
Current weights, four years of issuer returns, timestamped event updates, market-factor returns, and a pre-outcome peer classification for 8 investable and 24 reference issuers.
A preview of the supplied event data
4 of 520 rows · event_updates.csv
| Issuer | Available to system | Earnings surprise | Guidance change |
|---|---|---|---|
| EQ01 | 31 Mar 2026 | -1.0 | +1.0 |
| EQ02 | 31 Mar 2026 | +1.0 | -1.0 |
| EQ03 | 31 Mar 2026 | -1.0 | +1.0 |
| EQ04 | 31 Mar 2026 | +1.0 | +1.0 |
The full packet also includes current portfolio weights, daily issuer returns, market-factor returns, and peer groups.
What only the investor could answer
None of those tables says how concentrated the investor is willing to be. Three rules needed to construct the portfolio are absent from the brief:
Better analysis cannot recover a client preference from market data. Before allocating, the system has to recognize the gap and decide whether to ask or assume.
The paths diverged here
Both systems completed the quantitative work. Only one paused to recover the investor's rules.
Asked, then acted
Itoflow · GPT-5.6-Sol high
0.550
Task score
Stopped before portfolio construction and asked one bundled question for the concentration limit, minimum holdings, and minimum funded weight.
Investor exchange
Asked: Should I use a conservative envelope based on the supplied portfolio, or will you provide the exact limits?
Client: Maximum weight per asset 30%; minimum holdings four; minimum funded weight 8%.
Received all three values, then submitted a feasible four-position allocation.
Submitted portfolio
- EQ01
- 0%
- EQ02
- 30%
- EQ03
- 30%
- EQ04
- 0%
- EQ05
- 30%
- EQ06
- 0%
- EQ07
- 10%
- EQ08
- 0%
Inferred, then acted
Codex · GPT-5.6-Sol high
0.187
Task score
Found the omitted limits but did not ask. It inferred a conservative range from the current portfolio and retained all eight positions.
Investor exchange
No question was sent to the client before submission.
Submitted a feasible but more conservative allocation without receiving the client limits.
Submitted portfolio
- EQ01
- 9.375%
- EQ02
- 15.625%
- EQ03
- 15.625%
- EQ04
- 15.625%
- EQ05
- 15.625%
- EQ06
- 9.375%
- EQ07
- 9.375%
- EQ08
- 9.375%
Both portfolios were feasible. Itoflow asked for the missing rules before allocating and scored 0.550. Codex inferred the rules and scored 0.187.
One example shows the mechanism. The full study asks whether the same advantage survives across different portfolio decisions and different kinds of missing investor information.
From one case to a controlled comparison
Same decision, two information conditions
The worked example captures one behavior. To isolate that behavior from general quantitative ability, every task is evaluated twice. The investment problem is unchanged; only the timing of investor-specific information differs.
Interactive
The multi-turn benchmark
The investor's rules are missing from the briefing. The system must recognize what it needs and ask before submitting an action.
All information supplied
The single-turn control
The same investor rules appear in the briefing from the start. The system can begin research immediately.
- 14
- tasks in the benchmark
- 3
- agent systems, 7 configurations
- 2
- information conditions
- 196
- completed evaluations
What the comparison revealed
Scores fell when systems had to recover missing information
With every investor constraint supplied in the initial briefing, the best Itoflow configuration reached a median of 0.778 and the best model-matched Codex configuration 0.774. All configurations scored lower when they had to identify and request the missing constraints.
Itoflow and Codex both run GPT-5.6-Sol, so that pair is the model-matched comparison; Claude Code runs Opus 5. Each system uses its own tools, code, and research process. The scorer judges only the submitted action.
Results as of 1 August 2026
Median score by information condition
Itoflow
GPT-5.6-Luna · XHigh reasoning
0.738 → 0.689change -0.048
Itoflow
GPT-5.6-Sol · Low reasoning
0.685 → 0.663change -0.023
Itoflow
GPT-5.6-Sol · High reasoning
0.778 → 0.489change -0.290
Codex
GPT-5.6-Sol · High reasoning
0.751 → 0.349change -0.402
Claude Code
Opus 5 · High reasoning
0.604 → 0.306change -0.297
Codex
GPT-5.6-Sol · Low reasoning
0.774 → 0.228change -0.546
Claude Code
Opus 5 · Low reasoning
0.492 → 0.107change -0.385
Feasible actions in the interactive condition
When a system assumes missing investor constraints, it may submit a portfolio that breaks a hard limit. Those actions remain in the results and receive a score of 0. With all information supplied, 96 of 98 submitted actions were feasible.
Itoflow
GPT-5.6-Luna · XHigh reasoning
14/14 feasible
Itoflow
GPT-5.6-Sol · Low reasoning
13/14 feasible
Itoflow
GPT-5.6-Sol · High reasoning
12/14 feasible
Claude Code
Opus 5 · High reasoning
10/14 feasible
Claude Code
Opus 5 · Low reasoning
10/14 feasible
Codex
GPT-5.6-Sol · High reasoning
8/14 feasible
Codex
GPT-5.6-Sol · Low reasoning
7/14 feasible
Interactive
The initial briefing omits investor constraints that cannot be inferred from market data. The system must ask for them before submitting an action.
Median is the primary summary because a few extreme task scores can move the mean. Feasibility is reported because a numerically attractive action does not count if it breaks the investor's rules.
| System | Configuration | Median | Mean | Tasks completed | Feasible actions |
|---|---|---|---|---|---|
| ItoflowGuardian: GPT-5.6-Luna xhigh | GPT-5.6-LunaXHigh | 0.689 | 0.562 | 14/14 | 14/14 |
| ItoflowGuardian: GPT-5.6-Sol high | GPT-5.6-SolLow | 0.663 | 0.575 | 14/14 | 13/14 |
| ItoflowGuardian: GPT-5.6-Sol high | GPT-5.6-SolHigh | 0.489 | 0.431 | 14/14 | 12/14 |
| Codex | GPT-5.6-SolHigh | 0.349 | 0.355 | 14/14 | 8/14 |
| Claude Code | Opus 5High | 0.306 | 0.331 | 14/14 | 10/14 |
| Codex | GPT-5.6-SolLow | 0.228 | 0.344 | 14/14 | 7/14 |
| Claude Code | Opus 5Low | 0.107 | 0.238 | 14/14 | 10/14 |
A score of 0 matches the task's sensible default, 1 is the strongest attainable action under the task model, and negative values fall below the default. Actions that break a hard investor rule also receive 0.
Itoflow rows include its Guardian, a built-in reviewer that checks the research and the final action before submission; the benchmark awards it no points. In the research paper, the interactive condition is the multi-turn benchmark and the all-information condition is the single-turn control.
Release 1.0 reports one completed evaluation per task and configuration. Multi-run confidence intervals are in progress and will be added to these charts.
Behind the medians
All 196 scores
The aggregate result is not uniform. This matrix shows where each system gained from complete information, where it remained robust without it, and where a submitted action broke an investor rule.
- How to read a score
- Negative is worse than the task's sensible default, such as leaving the existing portfolio unchanged. 0 is no economic improvement over that default. Values toward 1 approach the strongest attainable decision under the task's own model.
- What × means
- The system completed the assignment, but its submitted action broke a hard investor constraint, such as a position limit. The action is preserved in the record and scored 0.
Swipe horizontally to compare configurations.
| Assignment | Interactive | All information supplied | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ItoflowHigh | ItoflowLow | ItoflowXHigh | CodexHigh | CodexLow | ClaudeHigh | ClaudeLow | ItoflowHigh | ItoflowLow | ItoflowXHigh | CodexHigh | CodexLow | ClaudeHigh | ClaudeLow | |
| Point-in-time corporate-event allocationSignals | 0.550 | 0.701 | 0.701 | 0.187 | 0.701 | 0.265 | 0.701 | 0.701 | 0.894 | 0.376 | 0.454 | |||
| Cross-market information-propagation allocationSignals | 0.025 | 0.000 | -0.366 | 0.538 | -0.010 | 0.082 | -0.582 | -0.582 | -0.582 | 0.288 | 0.288 | 0.196 | 0.000 | |
| Tactical allocation under uncertain market conditionsSignals | 0.000 | 0.192 | 0.132 | 0.337 | 0.212 | 0.195 | -0.105 | 0.212 | 0.116 | 0.038 | ||||
| Market-neutral equity allocationPortfolio construction | 0.915 | 0.899 | 0.907 | 0.852 | 0.862 | 0.830 | 0.748 | 0.856 | -0.051 | 0.907 | 0.834 | 0.907 | 0.865 | 0.859 |
| Drawdown-aware strategic allocationPortfolio construction | 0.783 | 0.856 | 0.954 | 0.731 | 0.457 | -0.009 | 0.953 | 0.912 | 0.953 | 0.953 | 0.973 | 0.666 | -0.202 | |
| Exclusion-aware regional completion portfolioPortfolio construction | 0.678 | 0.176 | 0.678 | 0.655 | 0.456 | 0.318 | 0.395 | 0.176 | 0.678 | 0.678 | 0.678 | 0.678 | 0.196 | 0.267 |
| Portfolio transition with competing research viewsPortfolio construction | 0.935 | 0.893 | 0.935 | 0.928 | 0.928 | 0.779 | 0.928 | 0.923 | 0.935 | 0.935 | 0.935 | 0.935 | 0.985 | 0.923 |
| Decide whether a market dislocation is investableRelative value | 0.000 | 0.257 | 0.320 | 0.536 | 0.645 | 0.337 | 0.343 | 0.545 | 0.664 | 0.402 | 0.649 | 0.347 | 0.541 | 0.529 |
| Schedule an approved transition over five sessionsExecution | 0.994 | 0.472 | 0.475 | 0.511 | 0.816 | 0.706 | 0.675 | 0.992 | 0.927 | 0.775 | 0.460 | 0.491 | 0.942 | 0.890 |
| Fund a defensive overlay for a concentrated sleeveRisk transfer | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.973 | ||||
| Protect an equity sleeve with a bounded-cost collarRisk transfer | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | |||||
| Build a liquid real-asset sleeve for persistent inflationRisk transfer | 0.733 | 0.870 | 1.000 | 0.733 | 0.733 | 0.870 | 0.870 | 0.870 | 0.870 | 0.870 | 0.914 | 0.965 | ||
| Replicate a target strategy with liquid proxiesInstitutional | 0.625 | 0.574 | 0.573 | 0.573 | 0.295 | 0.693 | 0.693 | 0.574 | 0.607 | 0.574 | 0.390 | |||
| Set an allocation from raw point-in-time extractsInstitutional | 0.427 | 0.301 | -0.306 | -0.689 | 0.453 | -0.614 | 0.908 | 0.800 | -0.944 | -0.270 | -0.269 | |||
Cell color shows score direction and magnitude; the numbers are the exact unclipped scores. A few hedging assignments sit near 1.000 for most configurations in the all-information condition. They chiefly test correct execution, and the methodology covers how much room each task leaves beyond it.
Inspect the underlying work
Open any assignment from request to score
The matrix is the summary. The task explorer below exposes what each system received, which investor rules changed between conditions, what action it had to submit, and how that action was judged.
Signals
Point-in-time corporate-event allocation
signal.event_response
Investor request
Improve an eight-equity sleeve after the latest earnings and guidance updates using only information available at the decision time.
Evidence shown
Current weights, four years of issuer returns, timestamped event updates, market-factor returns, and a pre-outcome peer classification for 8 investable and 24 reference issuers.
Required action and economic objective
Eight long-only opening weights.
Expected incremental terminal return versus the supplied portfolio, net of opening turnover cost.
Observed approaches on this task
Itoflow · GPT-5.6-Sol high
0.550Stopped before portfolio construction and asked one bundled question for the concentration limit, minimum holdings, and minimum funded weight.
Investor exchange
Asked: Should I use a conservative envelope based on the supplied portfolio, or will you provide the exact limits?
Client: Maximum weight per asset 30%; minimum holdings four; minimum funded weight 8%.
Received all three values, then submitted a feasible four-position allocation.
Feasible actionCodex · GPT-5.6-Sol high
0.187Found the omitted limits but did not ask. It inferred a conservative range from the current portfolio and retained all eight positions.
Investor exchange
No question was sent to the client before submission.
Submitted a feasible but more conservative allocation without receiving the client limits.
Feasible actionApplication to Itoflow
Why missing investor constraints matter
Investors do not always provide every constraint upfront. Itoflow is built to ask for decision-relevant values before proposing a portfolio. BuySideBench tests that behavior directly.
Each result scores the submitted action, not Itoflow's internal workflow, evidence system, or review machinery.
Run your system against BuySideBench
The next result on this page should not have to come from us. Model developers, researchers, and investment firms can request access while the public release is prepared.
The public repository will include the assignments, harness and scorer interfaces, and the records behind the results on this page. The research paper will document task construction and evaluation in full. Until those materials are published, we are onboarding participants directly.
- Public GitHub
- Coming soon
- Research paper
- Coming soon
- Early access
- hello@itoflow.ai