TransitFix All articles
Transit Technology

Why the Arrival Time on Your Transit App Is Often Wrong—and What Cities Are Doing About It

TransitFix
Why the Arrival Time on Your Transit App Is Often Wrong—and What Cities Are Doing About It

Photo: User:B, CC BY-SA 4.0, via Wikimedia Commons

You are standing on a train platform. Your phone shows the next southbound train arriving in four minutes. You wait. Four minutes pass. The app now says six minutes. You wait again. The train arrives eleven minutes after your original estimate, with no explanation offered by the app, the agency, or the overhead display board, which showed a different time still.

If this experience feels familiar, you are not alone—and you are not imagining things. Transit arrival data across a significant share of American cities is measurably, demonstrably inaccurate, and the gap between what apps promise and what buses and trains actually deliver has become one of the most corrosive forces eroding public trust in urban transportation.

This is an investigation into why that gap exists, which cities are leading the effort to close it, and what riders can do in the meantime.

The Anatomy of a Bad Data Feed

To understand why real-time transit information so frequently misleads riders, it helps to trace the journey that data takes from a moving vehicle to a smartphone screen.

Most transit agencies in the United States rely on a combination of GPS transponders installed on vehicles, automated vehicle location (AVL) systems that process and transmit positional data, and schedule-based prediction algorithms that estimate future arrivals based on a vehicle's current position and historical performance. The output of this system is typically formatted in a standard called GTFS-Realtime—a specification developed by Google and now maintained as an open standard—which third-party app developers consume to power the countdown timers riders depend on.

At each step in this chain, things can go wrong. GPS transponders malfunction or fall out of calibration. AVL systems transmit data at intervals that may be 30, 60, or even 90 seconds apart—an eternity when a bus is moving through urban traffic. Prediction algorithms that rely on historical averages perform poorly during disruptions, special events, or unusual weather. And the APIs through which third-party developers access agency data are sometimes throttled, cached, or simply not updated frequently enough to reflect ground truth.

"The term 'real-time' in transit data is often aspirational rather than literal," says Dr. Jonathan Reeves, a transportation data researcher at the Georgia Institute of Technology. "What agencies are frequently delivering is near-real-time data with meaningful latency. Whether that latency matters depends on how fast the vehicle is moving and how crowded the platform is."

The Infrastructure Deficit

The root cause of many data accuracy failures is not software—it is hardware. A substantial portion of the American transit fleet is operating with AVL equipment that was installed a decade or more ago, designed to meet the technical standards of its era rather than the expectations of contemporary riders.

The Federal Transit Administration's most recent National Transit Database figures indicate that approximately 30 percent of heavy rail vehicles and nearly 40 percent of bus fleets nationally are operating beyond their useful life benchmarks. Aging vehicles are more likely to carry malfunctioning tracking equipment, and the cost of retrofitting older rolling stock with modern sensors often competes with more pressing capital needs.

Beyond the vehicles themselves, the backend infrastructure that processes and publishes location data is frequently outdated. Some agencies are running AVL management software that has not received a major update in years, built on databases that were not designed to handle the volume and velocity of data that modern real-time systems require.

In a 2023 audit of 15 major metropolitan transit agencies conducted by the Eno Center for Transportation, researchers found that only six agencies could demonstrate that their published real-time feeds were updated at intervals of 15 seconds or less—the threshold generally considered necessary for meaningful accuracy in dense urban environments. The remaining nine agencies had update intervals ranging from 30 seconds to more than two minutes.

When the App Lies: Three Failure Modes

Riders who have experienced the frustration of inaccurate transit apps are typically encountering one of three distinct failure modes, each with different causes and different implications.

Ghost vehicles are perhaps the most disorienting. These are buses or trains that appear on a map or countdown display but do not actually exist in the location shown—typically because a vehicle's GPS unit has stopped transmitting and the system is projecting its expected position based on the schedule. The vehicle may be sitting in a garage, stuck in traffic three miles back, or operating on a detour. The app has no way to know, and so it shows riders a confident fiction.

Schedule bleed occurs when an agency's real-time system defaults to schedule-based predictions rather than actual location data, typically because the live feed has failed or become unreliable. In this mode, the app essentially shows riders when the bus is supposed to arrive, not when it will actually arrive. During periods of traffic disruption or service irregularity—precisely when accurate information is most valuable—the app becomes least reliable.

Data pipeline delays represent a subtler but pervasive problem. Even when GPS data is accurate and being transmitted correctly, it must travel through multiple processing layers before appearing on a rider's screen. Each layer introduces latency. By the time a commuter reads "3 minutes" on their phone, the underlying data may already be 45 seconds old—a meaningful discrepancy when the bus in question is traveling at 25 miles per hour through an intersection.

Cities Leading on Data Transparency

Not all transit agencies are struggling equally. A cohort of cities has made measurable progress on data accuracy, and their approaches offer a useful roadmap for others.

New York City Transit, which operates the largest bus fleet in North America, undertook a comprehensive AVL overhaul beginning in 2019 as part of its Bus Network Redesign initiative. New transponders were installed across the entire bus fleet, and the agency migrated to a new data platform capable of publishing location updates every 15 to 30 seconds. Independent audits conducted by TransitCenter in 2022 found that MTA Bus Time accuracy had improved substantially, with ghost vehicle incidents declining by approximately 60 percent compared to 2018 baseline measurements.

Chicago's Regional Transportation Authority has invested in a unified data platform that aggregates feeds from CTA buses, CTA rail, and Metra commuter rail into a single standardized output. While the underlying accuracy of each mode still varies, the consolidation has reduced the number of data handoffs where errors accumulate, and the agency has published a public-facing data quality dashboard that allows both developers and riders to monitor feed health in real time.

Seattle's King County Metro has implemented what it calls a "data stewardship" framework, in which a dedicated team is responsible not just for publishing data but for actively monitoring its accuracy and flagging anomalies for investigation. The agency also maintains an open dialogue with third-party app developers through a formal feedback channel, allowing developers who detect data anomalies to report them directly to the agency's technical team.

The common thread across these examples is intentionality. Improved data accuracy does not happen as a byproduct of other investments—it requires dedicated resources, clear accountability, and a willingness to measure performance honestly.

The Vendor Accountability Gap

One underappreciated dimension of the data accuracy problem involves the contractual relationships between transit agencies and the technology vendors who supply their AVL and data management systems. In many cases, vendor contracts specify that systems must be "operational" without defining accuracy thresholds or update frequency requirements in measurable terms. This creates a situation in which a vendor can deliver a technically compliant product that nonetheless produces unreliable rider-facing data.

"Agencies need to get much more specific in their procurements," argues Webb, the transit planner. "If you want 15-second update intervals and less than two percent ghost vehicle occurrence, write that into the contract. Make it a performance standard with consequences."

Some agencies are beginning to do exactly that. The Los Angeles County Metropolitan Transportation Authority's most recent AVL procurement included explicit data quality benchmarks with financial penalties for sustained non-compliance—an approach that transit policy advocates have urged other large agencies to adopt.

What Riders Can Do Right Now

While systemic improvements work their way through procurement cycles and capital budgets, riders are not entirely without recourse. A few practical strategies can improve the reliability of the information commuters act on.

First, prefer apps that display raw vehicle location data on a map rather than only showing countdown timers. Seeing where a bus actually is on a street grid allows riders to make independent judgments rather than trusting a prediction algorithm. Transit, Citymapper, and—where available—agency-native apps with map views all offer this capability.

Second, treat countdown timers showing more than eight minutes with appropriate skepticism. Prediction accuracy generally degrades as the time horizon extends, and a bus showing 12 minutes away is far more likely to surprise you than one showing 90 seconds.

Third, when you encounter clear data errors—ghost vehicles, buses that never appear, arrival times wildly inconsistent with reality—report them. Most agencies have feedback mechanisms, and some, like King County Metro, actively use rider reports to identify systemic data problems.

A Solvable Problem

The accuracy crisis in transit data is not a technological inevitability. The tools to deliver genuinely real-time information exist, have been demonstrated at scale, and are being deployed successfully in a growing number of American cities. What has historically stood in the way is a combination of underinvestment, diffuse accountability, and a tendency to treat data quality as a secondary concern relative to vehicle operations.

That framing is changing. As transit agencies compete for ridership against private mobility alternatives that offer precise, reliable pickup windows, the cost of inaccurate information has become impossible to ignore. The countdown timer on your phone may still lie to you today. With sustained pressure and investment, it does not have to tomorrow.

All Articles

Related Articles

The Final Stretch: How America's Transit Systems Are Losing Riders in the Last Half-Mile

The Final Stretch: How America's Transit Systems Are Losing Riders in the Last Half-Mile