School Decision Data Platform

Our Methodology

A note on originality: School Decision's data corpus is constructed from raw federal and state public records, but the published product is a new work of synthesis, comparison, and editorial interpretation.

A New Work of Synthesis

Government agencies publish raw tables; we publish a contextualized, comparison-derived analytical product. Every profile in our corpus integrates information from approximately a dozen distinct source datasets, cross-referenced against geographic, economic, and demographic baselines, then rendered as original editorial content.

The transformation from input to output is non-trivial and represents the substantive analytical work of the platform.

Source Material as Input

Raw inputs are sourced exclusively from authoritative federal and state datasets: the U.S. Department of Education's Common Core of Data, the Civil Rights Data Collection, the EDGE program's geocoded school-locator file, Small Area Income and Poverty Estimates, state-level assessment exports from individual state education agencies, and macroeconomic indicators from FRED and Zillow research.

These are public-domain records, but they exist in heterogeneous formats such as fixed-width files, pipe-delimited exports, state-specific encoding conventions, and inconsistent schema across federal and state boundaries. No source publishes them in a form usable by a parent making a school-selection decision. The work of making them usable is itself the value.

Cross-Source Synthesis

Our analytical layer integrates approximately twelve distinct source datasets per school profile, resolving identifier conflicts across federal and state systems and building the cross-reference machinery (NCES-to-FIPS crosswalks, EDGE geocoded coordinates, charter-status resolution) that allows a single school's record to inhabit a unified analytical frame. This integration step is not present in any source dataset.

A parent looking up Adacao Elementary in Guam, or a school in rural West Texas, or a Brooklyn middle school, sees a profile that synthesizes academic performance against state and county baselines, demographic context drawn from federal civil rights data, housing market context from Zillow, and economic context from FRED. None of which the source agencies present together. The synthesis is original work.

Comparative Analysis

For each school and district, we compute peer-comparison artifacts that situate the entity within county, metro, and state contexts. These comparisons including academic performance percentiles relative to similar districts, demographic composition relative to county baselines, and financial metrics relative to state norms are computed by our analytical pipeline against the unified corpus.

No source agency publishes these comparisons. They are derivative analytical works produced by School Decision and are protected as such.

Editorial Narrative Authorship

The most substantive layer of original work is the editorial narrative, a paragraph-form, parent-actionable interpretation of each school's and district's positioning. These narratives are not summaries of source data. They are original editorial works that synthesize multiple data sources, apply consistent analytical framing across the entire corpus, and surface what matters for a family making a school decision.

The narratives are produced through a proprietary three-stage research-and-writing pipeline:

Stage 1: Research

Independent contextual research is performed against open web sources for each locale and district, surfacing parent-relevant developments like recent policy changes, capital investments, demographic shifts, and district initiatives that contextualize the analytical data. These context briefs are themselves original derivative works, drawn from media coverage and authoritative reporting, synthesized into a structured format suitable for downstream analysis.

Stage 2: Comparative Analytics

Deterministic peer-comparison artifacts are generated from the unified corpus, situating each entity within geographic and demographic peer cohorts.

Stage 3: Editorial Synthesis

A constrained editorial writing step assembles the contextual research, comparative analytics, school-level facts, and district programs into a parent-actionable narrative. The narrative follows School Decision's editorial standard: factually grounded, action-oriented, free of decontextualized rankings, and respectful of the limitations of the underlying data. Narratives explicitly omit specific figures that would reduce the editorial work to a recitation of source-data tables; the prose conveys interpretation, not transcription.

Editorial Standard

School Decision applies a consistent editorial standard across the entire corpus that no source agency applies to its own data. We do not publish proprietary ratings. We do not aggregate test scores into single-number quality scores. We do not present federal civil rights data in ways that would invite reductive judgments about schools.

We surface inactive or alternative-program facilities with appropriate context rather than burying them. We attribute every factual claim to its source. This editorial frame is the platform's signature voice and is consistent across approximately 100,000 schools and 19,000 districts.

School Decision

Data Quality and Integrity

Our pipeline implements production-grade quality controls that source agencies do not apply to their published data. Every record carries immutable lineage to its originating source, version, and retrieval date. Every artifact is content-addressed with SHA-256 verification. Every state's coverage is validated against expected entity counts before deployment.

State-specific structural quirks such as Texas's Education Service Centers, New York's BOCES, and California's County Offices of Education are systematically identified and excluded from consumer-facing surfaces through a shared analytical filter rather than handled case-by-case. The corpus represents not just the integration of source data, but the application of consistent quality standards across heterogeneous federal and state sources.

Scale and Throughput

The platform currently maintains profiles for approximately 100,000 K-12 public schools, 19,000 districts, and an additional 4,000 international schools across 181 countries. The editorial narrative layer covers 34,000+ entities across four U.S. states including Florida, New York, Texas, and California representing approximately 30% of the U.S. K-12 school population.

Each narrative is an original editorial work; the cumulative corpus represents tens of millions of words of original synthesis. The platform's per-narrative production cost has been optimized to approximately $0.012, reflecting a substantial investment in pipeline efficiency that allows School Decision to apply consistent editorial treatment to a corpus an order of magnitude larger than any human-edited education-information platform.

Republication Policy

Source data is republished in its original form with attribution. Synthesized comparisons, analytical derivations, and editorial narratives are original works owned by School Decision and are not licensed for redistribution.

Researchers interested in source data should consult the originating agencies. Researchers interested in our analytical or editorial work should contact School Decision directly.