Since Data.gov launched in 2009 with 47 datasets and grew to hundreds of thousands of catalog entries, and since states and cities followed with portals of their own, the open-data question has shifted from whether to whether the data is usable. The U.S. policy spine is the OPEN Government Data Act, enacted in 2018 as part of the Foundations for Evidence-Based Policymaking Act: it requires federal agencies to publish machine-readable open data by default, favor open licenses and standard formats, and coordinate through the Data.gov catalog. Compliance with publication, annual reports on the statute show, is broad. Usability is narrower — and the difference is the story.
The four properties that make a dataset an asset
Schema stability: fields, codes, and meanings stay fixed across releases, or change with documented versioning. Nothing kills civic use faster than a column that quietly changes meaning mid-year. Stewardship: a named owner who fixes errors, answers questions, and keeps the refresh cadence the portal claims. Licensing: a machine-readable open license — the statute directs open licenses for federal data — so downstream users know reuse is legal. And a real consumer: data published because someone uses it, not because a dashboard counts it. Portals that score well on all four are rare and obvious: weather and GPS data, transit feeds in the General Transit Feed Specification, census tables — datasets with constituencies that complain loudly when quality slips.
How portals go wrong
The failure patterns repeat across the thousands of city, state, and federal portals. Trophy metrics: success measured in datasets published rather than downstream use, which rewards splitting one useful table into twenty thin extracts. Orphaned data: the employee who maintained the extract leaves, the refresh quietly stops, and the portal keeps serving stale numbers with no tombstone. PDFs wearing data clothing: machine-readable in name, with values locked in formatted reports. And licensing silence — no stated terms, which for cautious institutional users is equivalent to unusable. The Government Accountability Office's reviews of the OPEN Government Data Act implementation have flagged exactly these gaps: inventories that exist on paper, metadata quality problems, and agencies unclear on which data assets carry confidentiality or security restrictions that legitimately limit release.
What good looks like in practice
The strong programs converge on the same disciplines. A data inventory maintained under the Evidence Act's data-management requirements, so agencies know what they hold before publishing it. Publication pipelines rather than manual uploads: the dataset refreshes from the system of record, which is why transit feeds stay current while one-off civic extracts rot. Demand-driven prioritization: portals that track which datasets are downloaded and requested, and respond — the behavior behind the durable success stories like transit and geospatial base layers. And a governance body — a chief data Officer with actual authority, which the Evidence Act also required agencies to appoint — that can adjudicate what publishes, in what form, under what license.
The accountability case, stated plainly
Open data earns its budget as accountability infrastructure, not decoration. Budget documents, contracts, inspection results, and service metrics published with stable schemas let journalists, auditors, and council offices check claims the source systems would otherwise bury. That is also the honest test of a portal: not how many datasets it lists, but whether the dataset that would embarrass the government is on it, current, and correctly licensed.
FAQ
What law governs U.S. open data?
The OPEN Government Data Act of 2018, part of the Foundations for Evidence-Based Policymaking Act — machine-readable open data by default, open licenses, and coordination through Data.gov.
How can I tell if a portal dataset is maintained?
Check the refresh history and change log, whether a steward is named, and whether the schema has stayed stable — orphaned datasets usually show silent staleness instead.
Why do some open datasets go stale?
Manual publication processes and departed maintainers: without a pipeline from the system of record, refreshes depend on someone remembering to upload.
For more context, read Open 311: How Cities Track Whether the Pothole Ever Got Fixed.
For more context, read open meetings law.
For more context, read government payment delivery.
