Insights · corpus rollup

What the MESSAI corpus says

Reporting completeness, calibrated prediction intervals, anomalous papers, discovered laws, learned causal edges, and cross-system transfers — sourced from the trainer’s most recent published artifacts. Each panel reports its own status; nothing is silently zeroed out.

Reporting completeness

23.1% corpus average · 289 scored papers · 10,824 indexed across 5 system types

Corpus average

23.1%

reporting completeness

MFC mean

8.7

params reported per paper

MEC mean

3.6

params reported per paper

SystemPapersElectrode
specs
Operating
conds
Electrical
meas
BioData
rep
MFC3,84772%58%45%35%32%
MES4,84268%55%42%32%28%
BES1,09665%48%38%28%25%
MEC96670%52%40%30%28%
MDC7355%38%28%20%15%

Source: ISMET 2026 abstract Fig 2A. Live corpus-rollup endpoint pending.

Calibration health

Awaiting trainer export

research/calibration.json not yet published by trainer

Run services/ml-engine/training/calibration.py to publish.

Prior-trust distribution

102 fitted parameters · median 20 papers · max 367

Calibrated

8

8% of params

Modeled

20

20% of params

Curated

71

70% of params

Flagged

3

3% of params

Papers per parameter

5-9
18
10-19
33
20-49
23
50-99
11
100+
17

Number of parameters with N supporting papers. Higher bins = stronger priors.

Flagged parameters (caveat the value)

8% of canonical parameters meet the calibrated threshold (PPC + LOO + convergence all pass). Paper-level reproducibility-score distribution (the 289-paper analysis in the abstract) is tracked as a post-launch trainer pass.

Anomalous papers

Run services/ml-engine/training/score_anomalies.py to publish.

Learned causal edges

1 candidate edge · HillClimbSearch + BICScore · 173 papers

  • powerDensitycurrentDensity

Discovered symbolic laws

2 successful fits · PySR (Cranmer 2023) — symbolic regression

showing top 2 by loss
  • powerDensityvs.currentDensityMFC · n=51

    (x0 + (0.22200376 / ((x0 + 2.3158035) * -1.8066067))) - 0.6148749

    Logan 2008 §3.4: P = V·I. At matched load V≈V_oc/2 ≈ const → P ∝ I

    loss = 0.2252 · complexity = 11 · R² ≈ 0.198

  • coulombic_efficiencyvs.cod_removalMFC · n=29

    ((x0 + -0.19921306) / ((x0 + (x0 + 0.46581972)) / 0.0033559396)) + -1.308836

    Sleutels 2012 §4: at high COD-removal, more substrate goes to biomass (not e-) → CE drops

    loss = 0.4429 · complexity = 13 · R² ≈ 0.210

Cross-system transfers

64 viable transfers · 26 within-system pairs analysed

ParameterSourceTargetn samples
powerDensityMFCMSC5
powerDensityMFCOTHER5
powerDensityMFCREVIEW5
powerDensityMSCMFC5
powerDensityMSCOTHER5
powerDensityMSCREVIEW5

58 more transfers not shown.

Generated 2026-05-10

About these numbers

Each panel reads a single JSON artifact under apps/web/public/data/computed/research/ at request time and rejects fabricated zeros — if an artifact is missing or below threshold, you see the honest status, not blank cards.

Artifacts are regenerated by the trainer (services/ml-engine/training/) on each release. See docs/ismet-2026-continuation.md for the full pipeline.