HotInfra '26 — co-located with ISCA '26 · Raleigh, NC

A benchmark for infrastructure agents.

Beyond Pass/Fail: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Twelve real operational incidents — from IPMI power recovery to silent data corruption — spanning hardware, local systems, distributed systems, and user applications. Beyond pass/fail: durable state, invariants, cleanup, and risk.

Task coverage by layer

L4 User Applications3 tasks
L3 Distributed Systems7 tasks
L2 Local Systems1 task
L1 Hardware1 task
12 tasks11 configs3 backends

01Leaderboard

Updated Jun 28, 2026

Mean effective score across all 12 tasks. Source: Table 3 in the paper.

#AgentModelMean score
1Claude CodeClaude Opus 4.7Anthropic90.1%
2CodexGPT-5.5OpenAI86.1%
3Claude CodeClaude Opus 4.8Anthropic85.7%
4Cursor CLIAuto* · routerCursor81.1%
5Gemini CLIGemini 3.5 FlashGoogle78.1%
6Cursor CLIComposer 2.5Cursor74.6%
7OpenCodeDeepSeek V4 FlashDeepSeek65.5%
8OpenCodeMiMo V2.5 ProXiaomi60.9%
9OpenCodeMiMo V2.5Xiaomi60.6%
10Claude CodeClaude Sonnet 4.6Anthropic59.2%
11OpenCodeDeepSeek V4 ProDeepSeek49.8%
OpenCodeGLM 5.2Zhipuin eval
Kimi CodeKimi 2.7 CodeMoonshotin eval

* Auto is Cursor's built-in model router, not a fixed checkpoint. Kimi 2.7 Code and GLM 5.2 evaluations are in progress.

02Per-task results

pass partial fail

Full 12 × 11 score matrix across published model configurations.

Task
Anthropic
Opus 4.7
Anthropic
Opus 4.8
OpenAI
GPT-5.5
Google
Gemini 3.5 Flash
Cursor
Auto*
Cursor
Composer 2.5
DeepSeek
DS V4 Flash
Anthropic
Sonnet 4.6
Xiaomi
MiMo V2.5 Pro
Xiaomi
MiMo V2.5
DeepSeek
DS V4 Pro
IPMI Power Recovery
Cassandra Dead Node
0/4
Cassandra NIC Split Brain
0.80
0/4
Cassandra Hung Recovery
Cassandra CORDS
Ceph Pool Degraded
7/8
7/8
7/8
7/8
7/8
4/8
7/8
3/8
0/8
Ceph Bootstrap
0.70
4/7
0/7
0.05
0/7
1/7
1/7
0/7
Fileserver RAID10
0/2
0/2
0/2
0/2
SLURM/Puppet Cascade
8/11
6/11
8/11
8/11
8/11
0/11
8/11
6/11
3/11
4/11
4/11
DB WAL Recovery
5/7
5/7
0/7
5/7
5/7
0/7
0/7
0/7
0/7
0/7
PgBouncer Drift
Pelican Key Mismatch
5/15
5/15
7/15
10/15
5/15
5/15
2/15
0/15
0/15
8/15
9/15

Fractions show verifier checks passed. Full methodology is in the paper.

03Failure patterns

Seven recurring failure modes. The gray field shows incidence across all models; overlay up to two configurations to compare exposure.

Destructive Dx
72.7
Hidden Config
41.4
Side Effects
5.9
Non-Persistent
11.8
Cleanup Miss
29.1
Fix Regression
9.1
Deploy Residue
36.4

Radius = severity when exposed (critical 100 · high 65 · medium 40). Field = incidence-weighted severity across all 11 models. Overlay up to 2 configurations.

04Task taxonomy

Contribute a task →

12 tasks across four infrastructure layers and 3 execution backends.

TaskLayerBackendDifficultyCharacteristics
IPMI Power RecoveryL1 HardwareCloudLab●○○Physical-level IPMI power cycling
Cassandra NIC Split BrainL2 Local SystemsCloudLab●●○NIC-caused network partition
Cassandra Dead Node RemovalL3 Distributed SystemsCloudLab●●○Dead node removal & ring repair
Cassandra Node Hung RecoveryL3 Distributed SystemsCloudLab●●○Hung node recovery via IPMI/network
Cassandra CORDS PropagationL3 Distributed SystemsVM Cluster●●●Silent data corruption; adapted from CORDS (FAST '17)
Ceph Pool DegradedL3 Distributed SystemsVM Cluster●●●Degraded pool repair & CRUSH fix
Ceph BootstrapL3 Distributed SystemsVM Cluster●●●End-to-end Ceph cluster deployment
Fileserver RAID10 Silent DiskL3 Distributed SystemsVM Cluster●●●RAID10 silent disk latency injection
SLURM/Puppet Config CascadeL3 Distributed SystemsVM Cluster●●●Unattended-upgrades + Puppet config cascade
DB WAL RecoveryL4 User ApplicationsContainer●●●Encrypted WAL recovery; adapted from Terminal-Bench 2.0
PostgreSQL/PgBouncer DriftL4 User ApplicationsVM Cluster●●●Long-term persistent config drift
Pelican Namespace/Key MismatchL4 User ApplicationsVM Cluster●●●Real-world CHTC incident: JWKS/namespace mismatch