Back to Socratic

Live corpus operations

Socratic coverage

Automatic live counts for the paper corpus and every active non-paper source.

Connecting to live metrics. Last attempt awaiting first update.

Live paper metrics

These values come from the running index, the current catalog sample, and the continuous full-text worker. They are operational measurements, not a manual status report.

Searchable papers
Waiting for live data

Current papers available in Socratic search.

Cataloged papers
Waiting for live data

Automatic sample of the live paper catalog.

Abstracts present
Waiting for live data

Automatic sample of abstracts present in the catalog.

Full-text acquisition
PDF queuedPDF savedPDF failedPMC bodies

Current PDF acquisition and PMC body-text parser checkpoints.

Catalog layers, shown to scale

The top edge is the live paper catalog. The lower edges show current abstracts and search availability at their actual proportions. Layer 4 tracks PMC full-text bodies written by the live parser, alongside the funnel because it comes from a separate processing stream. Every ETA is tied to a declared corpus target. Layer 2 is a current catalog snapshot, not a finite abstract acquisition target: future source refreshes or new feeds can increase it. Search and PMC body text forecast from the active worker rate.

The proportional view will appear as soon as all three live catalog layers respond.

Search versus catalog

Searchable papers are ready to find now. Catalog and abstract values are refreshed automatically from the broader OpenAlex corpus while the search materializer keeps bringing those records into the user-facing index.

Full-text work

The PDF queue, saved, and failed counts come directly from the active acquisition worker. The fourth layer reports confirmed PMC body-text writes, so both acquisition and usable body-text progress stay visible.

Unified graph release

This is the graph publication state, separate from source availability. A source can be loaded before its claims are normalized and published into a pinned graph release.

The graph release service has not responded yet. Source metrics above remain independent and continue refreshing.

Non-paper sources

Every non-paper source with a Socratic loader is listed below. Each count is for that source's main loaded dataset, so the rows are not added together into one misleading total.

Connecting to the live non-paper source inventory.

Sources
Waiting for live data

Distinct non-paper source datasets in the live inventory.

Live
Waiting for live data

Sources whose main dataset is currently available.

Loading now
Waiting for live data

Sources with a loader actively refreshing their main dataset.

The complete source table will appear as soon as the live source inventory responds.

Paper-like universe by type

This is the declared scope reference used to describe the initial paper universe. The live metrics above report what is currently in the operating corpus.

Paper typeWorks
Journal articles197,096,571
Book chapters22,012,800
Conference papers16,142,337
Dissertations11,116,997
Preprints7,370,466
Reports2,268,074
Reviews831,190
Data and software papers30,929
Total paper-like universe256,869,364