Key findings
- OpenPhil is the dominant artery; peak 2021 ($81.7M); 2017 inflated by a single $30M OpenAI grant.
- Governance funding exploded 2018→2023 ($0.4M→$18.4M) — by 2023 the biggest OpenPhil subtrack after technical_safety ($24.6M).
- SFF grew explosively: 2020 $5.4M → 2023 $42.3M → 2025 $34.9M.
- FTX collapse (Nov 2022) was a funding shock (Redwood $6.6M, Ought $5M and other regrants cut).
- interpretability publications exploded (denoised 69→657, 2021–25); RLHF kept rising (25→1777, 2021–25) = absorption into mainstream.
- interpretability got a 2024-25 second wind (Scaling Monosemanticity 2024; circuit-tracing 2025); ai_control born 2023 (Redwood) → agentic control evals by 2025 (Ctrl-Z).
- Government money overtook philanthropy: from 2023 half a dozen national institutes (UK AISI ~$159M, ARIA $74M, Canada CAISI $36.5M, NSF $20M, Australia $19.7M) turn 'one fragile OpenPhil artery' into a broad, increasingly governmental flow.
- A new separate lens — VC/equity into safety startups (~$268M: Protect AI $108.5M, HiddenLayer/Goodfire $50M each, Gray Swan $40M, Lakera $20M); Goodfire's $50M rivals OpenPhil's whole annual technical budget. Never summed with grants.
- Money and attention diverge: interpretability attention ~×36.9 ahead of its ~$1M grants; governance / scalable_oversight / agent_foundations have money ahead of publications.
- evals caught up overnight: $0 → ~$96.7M in a single 2024-25 burst (AI Safety Fund, US AISI/NIST, Schmidt + international institutes).
- field_building is the largest single money flow ($326.1M itemized).
- 2024 was a record founding year (24 orgs): 'consolidation' is really proliferation into narrow interpretability / AI-security / CoT-monitoring startups.
- Ending arc: in 2025 both institutes dropped 'Safety' (UK → AI Security Institute, US → CAISI); FHI closed (2024); OpenAI Superalignment dissolved (2024); OpenPhil renamed itself Coefficient Giving (Nov 2025).
Caveats
- arXiv keyword counts are a proxy, not bibliometrics (term lags the real track start); 'AI control'/'dangerous capabilities'/'value learning' are especially noisy (flagged per row).
- Event count per track = collection density, not real-world activity — use money + publications; recent years are under-sampled (shaded on the emergence chart).
- Four money lenses that never sum: itemized grants (grant_out, $763.6M), donor yearly totals (funding_total, e.g. OpenPhil ~$304.5M), VC/equity (~$268M, expects a return), and pledges/budgets/estimates (context only, non-additive). The same fund reads differently per lens (OpenPhil $176.8M itemized vs ~$304.5M donor total). Per-track money views use grant_out only.
- Money mixes types — philanthropy, government (now international) and corporate, i.e. 'different dollars'; some are stated program budgets (UK £100M, ARIA £59M), not disbursements (currencies converted at reference rates).
- Per-track arXiv proxies overlap (umbrella/nested tracks) and can't be naively summed; the honest field-level line is the dedup _safety_corpus.
- OpenPhil 2024 total ($28M) is a different methodology than vipulnaik (org self-reports ~$50M); 2026 money is not yet available (honest gap, not zero).
- Radar uses ONE money axis only — multi-metric radars mislead on mixed scales, so we don't build them.
- Causal 'why' claims are interpretation over verified events, not sourced facts.
Open collection debts (flagged, not invented)
- arXiv curves for the 5 newer tracks are raw exact-phrase proxies (noise flagged); a citation-graph count would be cleaner but is out of scope.
- 2026 track-classified money is not yet machine-readable — left as an honest gap, not fabricated.
- Recipient-organisation is still free text in `detail` (not a structured column), so funder→org flows remain out of reach.
- SFF grants are annual totals, not track-classified, so SFF is absent from per-track money views (documented gap).
- Several government figures are stated program budgets, not disbursed amounts (flagged in the data).
Highlights
Timeline
A diagonal from bottom-left to top-right — early macrostrategy/agent-foundations give way to interpretability, evals and AI control; the bottom panel shows the same activity-mix shift.
Every event by organisation & year; track colour is a chronological gradient (ordered by first appearance). Bottom panel: stacked-area of events/track/year showing the activity-mix shift.
How to read. Each dot is one event; the org axis is sorted by first appearance, so dots form a diagonal from bottom-left to top-right. Track colour is a chronological gradient by birth year; stars mark milestone papers, eras are hatched. The bottom panel stacks events/track/year in the same colours.
Question it answers. Did the field move in one direction over time, and from what to what?
What you see. A clean diagonal: early macrostrategy and agent-foundations give way to interpretability, evals and AI control; the bottom panel shows the same activity-mix drift.
Look here. Follow the colour drift left-to-right and the rising slope of the bottom stack; ignore absolute heights (the bottom panel is collection density, not real-world activity).
How the field shifted across eras
Across the eras: 'foundations & strategy' falls 67% to 6% while 'technical safety' rises to ~60%; government goes from nothing to 24% and industry to 26%; grant money explodes from ~$0 to ~$335M with a dominant grey government band by 2024-26.
Cross-era comparison (era on X): family mix (100%), actor mix (100%), money by funder ($), org births/closures.
How to read. Four comparable panels, five eras on the X axis. Panels 1–2 are 100%-stacked shares (family mix; actor mix), panel 3 is absolute grant $ by funder, panel 4 is org births (up) vs closures/pivots (down). `n` above the columns is the era's total events, so normalising doesn't hide ~10x growth.
Question it answers. Across the whole history, who and what took over the field, and did the money follow?
What you see. 'Foundations & strategy' falls 67%→6% while 'technical safety' climbs to ~60%; government goes 0→24% and industry to 26%; grant money explodes ~$0→~$335M with a dominant grey government band by 2024–26.
Look here. Read the four panels top-to-bottom for one era column at a time — the shrinking blue, growing orange, arriving government, and swelling money bars all line up.
Cumulative money by funder type
Through 2020 it is almost entirely one blue philanthropy artery; from 2023 the orange government band explodes as national institutes arrive. The dotted green VC line (~$268M) is a separate lens, never summed.
Cumulative $ over time, grants stacked by type (philanthropy/government/corporate) — the money explosion & government takeover from 2023; VC/equity as a dotted separate-lens line, never summed.
How to read. Cumulative grant dollars over time, stacked by funder type (philanthropy blue, government orange, corporate red). The dotted green line is VC/equity (~$268M) laid over the same axis but never added into the stack.
Question it answers. How did the field's money base change — one artery, or many?
What you see. Through 2020 it is almost entirely one blue philanthropy artery; from 2023 the orange government band explodes as national institutes arrive at once.
Look here. Find where the orange band suddenly widens (2023) — that is the state entering; keep the dotted green line mentally separate, it is a different kind of money.
The shift over time
Track lifespans (dumbbell)
Old is not dead: reward modeling's last event is 2017 but RLHF publishes to 2025; agent foundations and value learning go quiet as events yet stay alive on arXiv — tracks get absorbed or move to the background.
Birth → last recorded event per track (ordered by birth). Hollow diamond = last year still publishing on arXiv (independent proxy). Old ≠ dead; absence of late events = collection density, not death.
How to read. Each line runs from a track's first recorded event to its last (tracks ordered by birth year). The hollow diamond marks the last year it still publishes on arXiv — an independent proxy, so the line and the diamond can diverge.
Question it answers. When a track stops producing events, is it dead — or just quiet?
What you see. Old is not dead: reward modeling's last event is 2017 yet RLHF publishes to 2025; agent foundations and value learning go quiet as events but stay alive on arXiv. Tracks get absorbed or move to the background rather than dying.
Look here. Compare the end of each line to its diamond — a big gap means the track lives on in publications after its last logged event.
Track emergence (event count)
Almost empty until 2012-13, explodes 2015-17 with the first technical tracks, then grows 2021-24 with evals, AI control and model organisms. Height is collection density, not real activity.
CAVEAT: collection density, not real activity; recent-years window shaded as incomplete.
How to read. Stacked-area of how many events of each track landed in the base per year. Height is collection density (how much was recorded), not real-world work; the recent, under-sampled years are shaded.
Question it answers. In what order did the tracks first appear, and when did they proliferate?
What you see. Almost empty until 2012–13; a burst in 2015–17 with the first technical tracks; then more growth 2021–24 as evals, AI control and model organisms pile on.
Look here. Read the order of first appearance and the shape of the bursts, not the absolute heights (and treat the shaded years as incomplete).
Organisational alluvial
The full route era to organisation to track. At first glance it is spaghetti (~115 orgs) — the next three variants roll it up to stay readable.
era → organisation → track. Ribbons show who worked on what, when (money rows excluded).
How to read. Three stacked columns — era (left) → organisation (middle) → track (right); each ribbon is a flow of events, its width the event count.
Question it answers. What is the full route from a period, through the org that acted, to the research track — for the whole field at once?
What you see. The complete route, but at ~115 orgs it is deliberately spaghetti: too many ribbons to trace individually. That is the point — it motivates the rolled-up variants that follow.
Look here. Don't chase single ribbons; note only the sheer density of the middle column, then move to the compact/by-era/pairs versions below for a readable read.
Alluvial compact (era → actor-type → family)
Rolled up to 7 actor types and 4 families: blue 'foundations & strategy' is dense early/left, orange 'technical safety' gains mass on the right.
Variant A — ~115 orgs collapsed to 7 actor types, 17 tracks to 4 families; wide legible ribbons.
How to read. Same three-column flow as the raw alluvial, but rolled up: era → 7 actor types → 4 track families (foundations & strategy, technical safety, governance & infrastructure, capabilities). Ribbon colour = family, width = event count.
Question it answers. Which family of work dominated, and how did that dominance shift over time?
What you see. A diagonal in colour: blue 'foundations & strategy' crowds the early eras / left, orange 'technical safety' gains mass toward the right.
Look here. Track where the blue vs orange ribbon mass sits along the era axis; the crossover is the field's centre-of-gravity shift.
Alluvial by era (small multiples)
One panel per era — 'foundations & strategy' rules 2005-2016, then the panels flood orange with 'technical safety' from the 2020s.
Variant B — one mini-Sankey per era: actor-type → track-family.
How to read. One mini-Sankey per era, each a 2-level flow actor-type → track-family; ribbon width = event count within that era.
Question it answers. Step by step, era by era, which family floods each period?
What you see. Blue 'foundations & strategy' rules 2005–2016; from the 2020s the panels flood orange with 'technical safety'.
Look here. Scan the panels left-to-right and watch the dominant ribbon colour flip from blue to orange.
Alluvial pairs (2-level Sankeys)
Two clean flows: the field's focus shifts to technical safety over time (left); researchers/industry pull technical safety while government and institutes pull governance (right).
Variant C — era→family and actor-type→family, two clean 2-level flows.
How to read. Two clean 2-level Sankeys side by side: left is era → family (the 'when'), right is actor-type → family (the 'who'); ribbon width = non-money events.
Question it answers. Separately: how did the field's focus shift over time, and which actors pull which family?
What you see. Left — 'technical safety' gains mass in the later eras. Right — researchers and industry pull technical safety, while government and policy/institutes pull governance & infrastructure.
Look here. On the right, follow the government and institute nodes: their ribbons run almost entirely into governance, not technical safety.
Event-type mix by era
Prehistory (2005-2012) is 100% org foundings; publications, grants and statements enter from 2013, and by the last era foundings are only ~28% of a far larger mix.
100%-stacked: how the activity mix shifted (founding → publishing → funding).
How to read. A 100%-stacked bar per era: the share of each event type (foundings, publications, grants, statements...). Heights are shares of the collection, not world volume.
Question it answers. What kind of activity was each era actually made of?
What you see. Prehistory (2005–2012) is 100% org foundings; publications, grants and statements enter from 2013; by the last era foundings are only ~28% of a far larger, richer mix.
Look here. Watch the founding slice shrink from full-bar to ~a quarter while publication and grant slices appear and grow.
The money
Money by funder, ranked
One bar per funder by type — philanthropy (OpenPhil $304.5M, SFF $144M...), government (UK AISI ~$159M, ARIA $74M...), corporate, and a separate VC/equity group (~$268M) never summed with the $763.6M of grants.
One horizontal bar per funder, sorted by canonical total, log axis, coloured by type (philanthropy/government/corporate) + VC/equity as a 4th group. Equity is a SEPARATE lens — never summed with grants.
How to read. One horizontal bar per funder, sorted by canonical total, on a log axis, coloured by type: philanthropy, government, corporate, and a distinct 4th group VC/equity. Log means each gridline is ×10, so bar lengths compress large differences.
Question it answers. Who put in the most, and what kind of money is it?
What you see. Philanthropy leads (OpenPhil $304.5M donor total, SFF $144M...), a wall of government institutes (UK AISI ~$159M, ARIA $74M...), corporate, and a separate ~$268M VC/equity group that is never summed with the $763.6M of grants.
Look here. Read the colour of each bar before its length — the point is the mix of money types, and the green VC bars share the axis only for scale, never a total.
Funding dot-strip (when & how big)
OpenPhil is active almost every year; the government cluster is unmistakable in 2023-2025 and VC rounds appear as green diamonds from 2023 — philanthropy carried the field alone until the state and market arrived together.
x=year, y=funder (sorted by total), dot area ~ that year's $, colour=type; VC/equity rounds as diamonds. Shows when each funder acted and the size of each move.
How to read. x = year, y = funder (sorted by total); each dot's area ≈ that year's dollars, colour = type; VC/equity rounds are diamonds. It shows timing and size that the cumulative and ranked views hide.
Question it answers. When did each funder act, and how big was each individual move?
What you see. OpenPhil is active almost every year; the government cluster is unmistakable in 2023–2025 (big orange dots); VC rounds appear as green diamonds from 2023.
Look here. Scan the 2023–2025 columns: philanthropy carried the field alone until the state (orange) and market (green diamonds) arrived together there.
Funding Sankey (funder → track)
OpenPhil and SFF spread across many tracks; government institutes enter with precision — mostly field-building, evals and robustness. Routes all itemized grants (~$764M).
Known individual grants routed from funder to research track.
How to read. A Sankey of itemized grants (~$764M): funders on the left, research tracks on the right, ribbon width = grant amount. Donor totals and VC/equity are excluded so nothing double-counts.
Question it answers. Who funds what — do funders spread across many tracks or specialise?
What you see. OpenPhil and SFF fan out across many tracks at once; the government institutes (UK/Canada/Australia/US AISI, EU AI Office, DARPA, NSF) enter with precision, mostly into field-building, evals and robustness.
Look here. Compare the number of ribbons leaving OpenPhil/SFF (many) versus a government institute (one or two thick ones) to see broad philanthropy vs targeted state money.
Funding routing (era → funder → track)
The same grants with an era level — the mass of money shifts into 2024-2026, and the new government institutes feed evals/field-building/robustness.
Same grants, 3 structural levels. Recipient-org is not a clean column, so it is not a level.
How to read. The same itemized grants as the previous Sankey, now with a third level on the left: era → funder → track. Ribbon width = grant amount.
Question it answers. When did the money flow, and which funders fed which tracks in each period?
What you see. The mass of money shifts into 2024–2026, and the new government institutes feed evals / field-building / robustness.
Look here. Follow the thick ribbons out of the 2024–2026 era node to see the government institutes as their dominant source.
Pledges, budgets & estimates (4th lens)
The 4th lens — money that is not a disbursed grant: launch pledges (FTX ~$160M), org budgets, overlapping field estimates, seeds. Heterogeneous and non-additive, so shown but never summed.
Collected money that is NOT a disbursed grant: launch pledges (FTX ~$160M pledged vs $18.7M disbursed), org annual budgets, overlapping field-wide estimates. Ranked bars grouped by sub-type on a log axis, NEVER summed with grants/donor totals/VC.
How to read. Ranked horizontal bars on a log axis, grouped by sub-type: launch pledges, org annual budgets, field-wide estimates, and seeds. These are heterogeneous and overlapping (estimates already contain budgets), so they are shown but never summed.
Question it answers. What money exists in the data that is NOT a disbursed grant?
What you see. FTX's ~$160M launch pledge towers over org budgets and overlapping field estimates — a 4th lens that keeps every collected dollar visible without faking a total.
Look here. Read magnitude only, by colour group; resist adding any bars together — that is the whole point of this lens.
FTX pledged vs disbursed (4th lens)
The sharpest gap: FTX pledged ~$160M at launch but disbursed only $18.7M before its Nov-2022 collapse — announced intent vs delivered money.
The 4th lens's headline as a dumbbell: FTX pledged ~$160M at launch but disbursed only $18.7M; the rest of the non-grant money (budgets, field estimates, seeds) ranked below. Never summed with grants/donor totals/VC.
How to read. A dumbbell for FTX — two dots, pledged (~$160M) vs disbursed ($18.7M), joined by a connector — with the rest of the non-grant money ranked below on the same log axis.
Question it answers. How far does announced intent diverge from money actually delivered?
What you see. The sharpest gap in the whole dataset: FTX pledged ~$160M at launch but disbursed only $18.7M before its Nov-2022 collapse.
Look here. Read the length of the FTX connector as the credibility gap between a headline pledge and a bank transfer.
Money vs attention & the fates of a track
Funding vs attention by track (small multiples)
One panel per track: interpretability is a steep attention line over tiny money bars (attention far ahead of money), evals is a wall of 2024-25 bars, AI control has money before its attention lifts.
One mini-panel per track: grant money/yr (bars, left axis) + arXiv attention (dotted line, right axis = log). All itemized grants ($763.6M) incl. a grey 'other (untracked/aggregate)' panel. Both are proxies.
How to read. Small multiples, one mini-panel per track (ordered by total grant money). Bars (left axis) = that track's itemized grant $/year; dotted line (right axis, log) = its arXiv attention. Each panel is legible on its own.
Question it answers. For each track, does money arrive before or after scientific attention?
What you see. Interpretability is a steep attention line over tiny money bars (~$1M grants) — attention far ahead of money; evals is a wall of 2024–25 bars (~$96.7M, nearly all in one year); AI control has money bars from 2021 before its attention lifts.
Look here. In each panel compare the bar heights to the dotted line's slope — a rising line over flat bars is 'science ahead of money'; tall bars under a flat line is the reverse.
Money vs attention (field)
Donor totals (OpenPhil+LTFF, bars) end at 2024 because 2025+ totals are not published yet; the solid red line is the trustworthy dedup corpus, the dotted line the inflated keyword sum.
Field-level: OpenPhil+LTFF yearly donor totals (bars, end at 2024 — 2025+ not published yet) vs arXiv attention (solid = dedup corpus, dotted = inflated keyword sum; right axis = log).
How to read. Four series: blue bars (left axis, $/yr) = OpenPhil+LTFF donor totals, de-duplicated; solid red line (right axis, log) = the trustworthy dedup `_safety_corpus`; dotted red line = the inflated keyword sum; grey band = the partial final year.
Question it answers. At the field level, do money and scientific attention track each other?
What you see. The bars stop at 2024 because donors haven't published 2025+ totals yet (a source gap, not 'money dried up'); attention keeps rising steeply — the two drift apart.
Look here. Trust the solid red line, not the dotted one — the gap between them is exactly how much the per-track keyword proxies overstate the real count.
Who leads per track (gap from field average)
Distance from the field-average money-to-attention rate: blue = attention ahead of money (interpretability ~x36.9), red = money ahead (scalable oversight, agent foundations, governance).
Diverging bars: per track log2(attention / field-average money-to-attention rate); blue = attention runs ahead of money (interpretability ×36.9), red = money ahead (scalable_oversight, agent_foundations, governance). Both are proxies.
How to read. Diverging bars: per track, log2(attention / field-average money→attention rate). Blue points right = attention ahead of money; red points left = money ahead. The ×N label is the factor away from the field average.
Question it answers. Which tracks are 'under-funded for their attention' and which are 'over-funded for their output'?
What you see. Interpretability is the extreme blue case (~×36.9 attention ahead of its ~$1M in grants); scalable oversight, agent foundations and governance sit on the red side, money ahead of publications.
Look here. Read direction (blue vs red) first, then the ×N magnitude — the longest blue bar is the sharpest 'science runs ahead of money' gap.
Money rank vs attention rank (dumbbell)
The same tracks by rank — money rank vs attention rank joined by a connector; interpretability's dots pull wide apart, well-matched tracks keep them close.
Companion to the gap chart: each track's rank by grant money (orange) vs rank by arXiv attention (green); long blue bar = attention far ahead of money, red = money far ahead, overlap = the two agree.
How to read. A dumbbell per track: its rank by grant money (orange dot) vs its rank by arXiv attention (green dot), joined by a connector (1 = smallest, N = biggest). This companion to the gap chart shows ordering rather than magnitude.
Question it answers. Does a track's money rank match its attention rank?
What you see. Interpretability's two dots pull wide apart; well-matched tracks keep their dots close, so the connector is short.
Look here. Scan for the longest connectors — a long blue one means attention ranks the track far above money; overlapping dots mean the two agree.
Track money footprint (radar)
Tracks' money footprint on a single log axis — deliberately one metric (money), since multi-metric radars mislead across scales.
Total grant $ per track on ONE log axis. Caveat: single-metric, not multi-axis.
How to read. A radar where every spoke is a track and the single radial axis is total grant $ on a log scale. It is deliberately one metric — multi-metric radars mislead across mixed scales, so none is built.
Question it answers. What is each track's relative money footprint at a glance?
What you see. An uneven star: field-building, robustness and evals reach far out, while interpretability and the small tracks sit near the centre.
Look here. Read spoke length as order-of-magnitude (log), not linear — a spoke twice as far out is ×10-ish more money, not double.
Organisations
Organisation lifespans (Gantt)
One bar per organisation, founding to last closure/pivot; 14 closure/pivot events collapse to 10 distinct orgs. Real closures (FHI, OpenAI Superalignment) land in 2024; FTX's 2022 collapse is the earlier exception.
Founding → closure/pivot (or ongoing). FHI/Superalignment closed 2024; institutes renamed 2025.
How to read. One bar per organisation, from founding year to its last closure/pivot (or 'still alive'), with a cross on that last event. Unlike the births/closures chart, this counts orgs, not events.
Question it answers. Which organisations actually ended or changed course, and when?
What you see. Real closures (FHI, OpenAI Superalignment) land in 2024; FTX's 2022 collapse is the earlier exception. The 14 closure/pivot events collapse to 10 distinct orgs here (MIRI, OpenAI, OpenPhil pivoted more than once).
Look here. Count the crosses (10 orgs), not the red events elsewhere — and note only 3 are true closures; the rest are pivots.
Births vs closures/pivots
New orgs up, closures/pivots down. Counts events, not orgs — 14 closure/pivot events (3 closures + 11 pivots), since one org can pivot several times (MIRI in 2013/2018/2024).
New orgs (up) vs closures + pivots (down) per year.
How to read. Bars per year: new organisations point up, closures and pivots point down. This counts events, not distinct orgs, so one org can appear several times.
Question it answers. Is the field expanding or contracting year by year?
What you see. Strongly up: births dominate; the 14 down-events (3 closures + 11 pivots) are outnumbered, and one org can pivot repeatedly (MIRI in 2013/2018/2024).
Look here. Don't equate a down-bar with a dead org — read it as 'change of focus'; cross-check the org-lifespan chart above for actual closures.