Live corpus operations
Socratic coverage
Automatic live counts for the paper corpus and every active non-paper source.
Connecting to live metrics. Last attempt awaiting first update.
Live paper metrics
These values come from the running index, the current catalog sample, and the continuous full-text worker. They are operational measurements, not a manual status report.
- Searchable papers
- Waiting for live data
- Cataloged papers
- Waiting for live data
- Abstracts present
- Waiting for live data
- Full-text acquisition
- —PDF queued—PDF saved—PDF failed—PMC bodies
Current papers available in Socratic search.
Automatic sample of the live paper catalog.
Automatic sample of abstracts present in the catalog.
Current PDF acquisition and PMC body-text parser checkpoints.
Catalog layers, shown to scale
The top edge is the live paper catalog. The lower edges show current abstracts and search availability at their actual proportions. Layer 4 tracks PMC full-text bodies written by the live parser, alongside the funnel because it comes from a separate processing stream. Every ETA is tied to a declared corpus target. Layer 2 is a current catalog snapshot, not a finite abstract acquisition target: future source refreshes or new feeds can increase it. Search and PMC body text forecast from the active worker rate.
Search versus catalog
Searchable papers are ready to find now. Catalog and abstract values are refreshed automatically from the broader OpenAlex corpus while the search materializer keeps bringing those records into the user-facing index.
Full-text work
The PDF queue, saved, and failed counts come directly from the active acquisition worker. The fourth layer reports confirmed PMC body-text writes, so both acquisition and usable body-text progress stay visible.
Unified graph release
This is the graph publication state, separate from source availability. A source can be loaded before its claims are normalized and published into a pinned graph release.
The graph release service has not responded yet. Source metrics above remain independent and continue refreshing.
Non-paper sources
Every non-paper source with a Socratic loader is listed below. Each count is for that source's main loaded dataset, so the rows are not added together into one misleading total.
Connecting to the live non-paper source inventory.
- Sources
- Waiting for live data
- Live
- Waiting for live data
- Loading now
- Waiting for live data
Distinct non-paper source datasets in the live inventory.
Sources whose main dataset is currently available.
Sources with a loader actively refreshing their main dataset.
Paper-like universe by type
This is the declared scope reference used to describe the initial paper universe. The live metrics above report what is currently in the operating corpus.
| Paper type | Works |
|---|---|
| Journal articles | 197,096,571 |
| Book chapters | 22,012,800 |
| Conference papers | 16,142,337 |
| Dissertations | 11,116,997 |
| Preprints | 7,370,466 |
| Reports | 2,268,074 |
| Reviews | 831,190 |
| Data and software papers | 30,929 |
| Total paper-like universe | 256,869,364 |