Stuck in the Past: Why America's Transit Data Infrastructure Is Failing Commuters
Photo: Torrance Transit, Public domain, via Wikimedia Commons
Imagine a city where every traffic light, every parking garage, and every commercial delivery fleet communicates in real time with a unified urban intelligence layer—adjusting signals, rerouting vehicles, and dynamically pricing curb access in response to live conditions. Now imagine that the city's own bus network is managed through a patchwork of spreadsheets, proprietary software contracts signed in the early 2000s, and scheduling databases that cannot talk to each other without manual export and re-import.
This is not a hypothetical. It is, to varying degrees, the operational reality of a significant share of American transit agencies today.
The Data Problem No One Talks About
Public conversations about transit technology tend to focus on the visible layer: the apps that display arrival times, the contactless fare readers, the real-time departure boards in station concourses. What receives far less attention is the foundational data infrastructure that sits beneath all of these consumer-facing tools—and which, in many systems, is profoundly inadequate.
At the core of the problem is fragmentation. A mid-sized American transit agency might operate its fixed-route bus network on one software platform, its paratransit service on a second, its fare collection on a third, and its maintenance scheduling on a fourth. These systems were often procured separately, from different vendors, over the course of decades. They were not designed to integrate, and in many cases they actively resist it. Data that exists in one system frequently cannot be accessed by another without expensive custom middleware or time-consuming manual processes.
The consequences are not merely administrative. When a bus breaks down and a supervisor needs to reassign vehicles, the lack of integrated data means that decisions are made with incomplete information. When a city planner wants to understand how transit ridership correlates with development patterns, the data may simply not exist in a form that supports that analysis. When a developer wants to build an application to help commuters navigate the system, they may encounter data exports formatted to standards that were obsolete before the iPhone existed.
The Legacy Vendor Problem
Understanding why these systems persist requires understanding the political economy of government technology procurement. Transit agencies are public entities, subject to procurement rules that favor competitive bidding, contract stability, and risk aversion. When an agency signs a long-term contract with a transit software vendor, switching costs become enormous over time: staff are trained on proprietary interfaces, workflows are built around system-specific limitations, and the institutional knowledge required to operate an alternative platform effectively atrophies.
Vendors, for their part, have historically had limited incentive to adopt open data standards or enable easy data portability. A transit agency whose operational data lives inside a proprietary platform is a captive customer at renewal time. Several transit technology vendors have faced criticism from agency staff and transit advocates for structuring contracts in ways that make data extraction costly and technically burdensome.
The situation is further complicated by the funding structure of American transit. Capital expenditures—vehicles, infrastructure, major technology systems—are often eligible for federal funding through programs administered by the Federal Transit Administration. Operating expenditures, including the ongoing costs of software maintenance and data management, are funded primarily through state and local sources that are frequently under pressure. Agencies facing budget constraints have limited appetite for the short-term costs of data modernization, even when the long-term benefits are clear.
What Open Data Actually Means
The term "open data" can encompass a range of practices, from publishing static schedule files in a publicly accessible format to streaming live vehicle location data through a real-time API. In the transit context, the most significant standard to emerge in the past two decades is the General Transit Feed Specification, or GTFS, originally developed through a collaboration between Google and Portland's TriMet in 2005.
GTFS and its real-time extension, GTFS-RT, have become the closest thing the transit industry has to a universal data language. When an agency publishes its schedule and vehicle location data in GTFS format, that data becomes immediately usable by a wide ecosystem of applications—Google Maps, Apple Maps, Transit App, Moovit, and hundreds of smaller tools developed by independent developers and civic technology organizations.
The impact of GTFS adoption has been measurable. Research published by transit technology researchers has found that agencies publishing open GTFS data see higher rates of third-party app development referencing their network, which correlates with increased rider awareness and, in some cases, ridership growth. The specification has also enabled researchers and policymakers to conduct comparative analysis across systems at a scale that was previously impossible.
Yet GTFS adoption remains uneven. Smaller agencies, particularly those serving rural and exurban areas, frequently lack the technical staff to generate and maintain accurate GTFS feeds. Some agencies publish feeds that are technically compliant but practically inaccurate—schedules that have not been updated to reflect service changes, or vehicle location data that lags reality by several minutes.
Cities Getting It Right
A handful of American cities have moved beyond basic GTFS compliance toward genuinely integrated open data ecosystems, and the results offer a compelling model.
The Los Angeles County Metropolitan Transportation Authority has invested substantially in its developer resources program, maintaining a well-documented API, publishing historical trip data in accessible formats, and engaging directly with the civic technology community through events and grant programs. The agency's open data posture has catalyzed a community of developers building accessibility tools, crowdsourced reliability trackers, and transit-integrated wayfinding applications that the agency itself could not have built or sustained internally.
In the Seattle area, King County Metro has been a national leader in integrating transit data with regional transportation management systems, enabling more sophisticated coordination between bus operations and traffic signal infrastructure. The agency's investment in data standardization has also made it a more attractive partner for mobility-as-a-service pilots, because potential partners can access reliable, machine-readable data without negotiating bespoke data-sharing agreements.
Chicago's Regional Transportation Authority has taken a different but equally instructive approach, focusing on making internal data more accessible to planners and analysts within the agency before externalizing it. The result has been more evidence-based service planning—a quieter but arguably more consequential benefit than consumer-facing applications.
The Developer Ecosystem Waiting to Be Unlocked
Perhaps the most underappreciated argument for transit open data is the innovation it enables outside agency walls. When transit data is accessible, accurate, and standardized, it becomes a platform—a foundation on which developers, researchers, and entrepreneurs can build tools that serve riders in ways no single agency could anticipate or fund.
Startups including Remix, now part of Via, have built sophisticated network planning tools on top of open transit data that agencies themselves use to model service changes. Civic technology organizations like OpenPlans have developed public-facing tools that translate complex network data into formats legible to community members participating in planning processes. Academic researchers have used open transit data to publish findings on equity, accessibility, and environmental impact that have directly influenced policy debates.
All of this activity depends on data being available. When it is not—when it lives in proprietary systems, formatted to obsolete standards, accessible only through expensive licensing arrangements—this entire ecosystem simply does not exist.
The Path Forward
Modernizing transit data infrastructure is not a glamorous undertaking. It does not produce ribbon-cutting moments or generate the kind of press coverage that a new light rail line or a flashy autonomous vehicle pilot attracts. But it is, in many respects, the prerequisite for nearly every other transit technology advancement.
Federal policy has begun to reflect this reality. The Infrastructure Investment and Jobs Act of 2021 included provisions encouraging transit data standardization and interoperability, and the FTA has signaled increasing interest in conditioning technology grants on open data compliance. These are meaningful shifts, even if implementation will take years.
For transit agencies, the practical first step is often less daunting than it appears: audit existing data assets, identify the highest-value datasets for external publication, and begin the process of generating accurate GTFS feeds if none exist. The technical barrier is real but manageable. The institutional barrier—the inertia of legacy contracts, the risk aversion of procurement culture, the absence of internal champions for data modernization—is frequently larger.
The commuters navigating American cities deserve transit systems that operate in the present tense. Getting there starts with the data.