CUSTOMER STORIES

Generating Inter-city Travel Insights From Mobile Location Data

Transcription

Transcript

This customer story has been adapted from a presentation given by Daniel Percy, Business Development Director at ATR.

Introduction: ATR and the Regional Aircraft Market

Hi everybody, I'm Daniel Percy. I work at a company called ATR — you may know one of our parent companies, Airbus. At ATR, we make regional aircraft, and we have a strong social focus: connecting communities. The people who use our aircraft are usually based in or live in isolated or small communities. In the British Isles, for example, our aircraft connect people who live in the Shetland Islands with cities in Scotland, or the Channel Islands with the mainland on the south coast. They often connect people across Ireland to Britain. What we do is very social.

The Challenge: The “Invisible” 97% of Ground Travel

We're interested in intercity journeys — travel between cities. In our market, the good news we found on our way in this data journey is that the vast majority of intercity journeys are regional, meaning between 100 and 400 nautical miles (around 440 US miles). In the four countries we've studied, the vast majority of intercity journeys by all modes are basically between 75 and 90%.

To put that into context, in the US market there are about 10 billion intercity journeys a year — a massive market, and roughly 75% of that is 7.5 billion journeys. You'd think that's good news for us and that there are lots of opportunities. But the tricky part is that the vast majority of these intercity journeys are actually going on the ground. In the US, of those roughly 12 billion journeys, 97% are actually going by ground — personal cars mostly in the US, but in countries like India a lot goes by train.

Our challenge was really to understand this competition and this massive market opportunity, for which there are no databases available. We can't go out and buy a database about personal journeys. If you jump in your car, it's not really recorded. Even bus or train journeys are very difficult to find data about. So 97% of our market opportunity was unquantified — something we needed to develop business intelligence about.

Introducing MobilityMonitor: Tracking 130 Million People

We created a tool we call MobilityMonitor. Normally, aircraft manufacturers talk about 20-year forecasts using data like Oxford Economics. But here, we're not using statistics to predict future demand based on air travel trends — we've used mobile location data coming from mobile phones and mobile GPS data to look at historical trends across many months in any given country. It's a gigantic volume of data: 130 million people, and we've so far mapped trends between a thousand airport catchments. This has been a research activity, and a few months ago we came out of the proof-of-concept phase.

The Data Science Journey: From Raw Data to Mobility Insights

We've gone through several core steps to turn raw data into real business intelligence. First was understanding the raw data itself — what mobile location data looks like in terms of quality, and how we can apply it to quantifying journeys. We had to transform billions of raw data points into millions of rows of trip information. We're not interested in trips everywhere — only those that could convert to air, from around one airport to another. So we had to learn how to do catchment analysis using this data and other features like isochrones, so we could create intercity origin-destination pair flows and measure mobility demand between cities.

Our goal is to see the full picture and understand what part of it could convert to air, to find new business opportunities for our airline customers. We learned to use existing logit modeling approaches to produce mode-share models. And, like with any data project, we had to fuse data — one data source on its own is never enough. We had to learn how to calibrate up from the section of the population captured in the raw data (which isn't the whole population) to the full population, in order to estimate total possible demand.

On raw data, there's a choice between telco data and mobile location data coming from apps and websites. We went with the apps/GPS location data mainly because those suppliers cover many different countries, and we're a global business — our customers are in 200 countries, so we needed a scalable solution. Telco data tends to be country-specific, and we didn't want a new supplier every time we wanted more data.

Telco data is super high quality — many points per device, near-perfect insight into position. Mobile GPS location data is much more variable. Looking at a month of data in Turkey, plotting points-per-device-per-day against number of unique H3 resolution-5 cells per device, you see a bit of everything: devices that produce huge numbers of points per day but don't move anywhere (not very useful), and devices that produce a low number of points but do move around, sometimes seen once a day in different locations (potentially useful). Only a small share of devices fall into the “not very useful” category — fortunately, a lot of devices across the rest of the spread enable us to produce trip information.

The Anatomy of an Air Trip vs. Ground Trip

In an ideal world, we'd have great-quality data: a clear trail of breadcrumbs showing exactly where someone started their journey, where they sat in the airport, how long they spent there — and the same on the other side, plus a clean, fast air segment (500+ mph) that's obviously distinguishable from any ground mode. In reality, it's a data science project, so a typical air trip looks messier: a good start, few breadcrumbs along the way, sometimes only one of two airports identified, and sometimes uncertainty about where a trip really ended. That affects the quality of the results — speed, distance, and time all get somewhat distorted.

So we had to understand our new trips dataset from its own data-quality point of view, with three core needs: identifying that somebody moved, having good air-mode identification so we can quantify the air-mode market, and using the data for trip time and distance models. We do use APIs for routing analysis to understand trip times, but we're talking about long journeys — in Indonesia, for example, journeys can go over water and take days even if the distance isn't huge, and we look beyond our regional market to quantify the full market. Where data quality is good enough, we compare API-based time and distance to the real time and distance we get out of the trips data.

Real-World Impact: Planning 100 New Airports in India

Once we had the trips, we had to build catchment analysis — and along the way we ended up creating things that are very relevant for airport development, not just airline and route development. This mobility and mobile GPS data turns out to be a really good source of information for developing countries building out their air networks. In India, the number of airports in operation serving scheduled flights has doubled over the last 10 years, and the government secured funding just last month to build another 100 airports over the next eight years.

That's a market challenge for us as an airframer, for the airlines, and for governments and authorities — figuring out where to put airports and what size the market will be where there's so much development and growth. It's very hard for businesses to project themselves into the future, and we're finding that using data based purely on observations of existing ground traffic is a powerful tool to bring insight into that process. In airport development, we can quantify mobility demand within a potential airport site, see how it competes with and overlaps with other potential sites, and help our airline customers choose between existing airports in a city — for example, comparing Gatwick and Heathrow.

Modeling Air Market Share and Route Demand

That's a bit of a byproduct — our main focus is connecting communities and helping our airlines find new air routes. The heart of the analysis is quantifying the existing market share of air at the city-pair level: our observations of what share air captures of total mobility between two cities. That part was relatively straightforward once we had the right data quality and the right selection of air trips.

What's genuinely new is that we normally know the total demand on an air route — there are plenty of aviation databases for that — but we didn't know that demand in proportion to the full mobility picture. We created this high-quality new dataset and partnered with Georgia Tech University, which had previously done this kind of analysis using bi-modal logit modeling to build market share from observations. In the past, the observations they used were theoretical, modeled from US government data; here, we could supply them with actual, true observations of air-mode markets.

The observations were the easy part. The hard part — the complex part of the model — is generating the other axis: a function of time and cost per mode, used to build these models. The result is a powerful trend along the expected S-curve these models generate, letting us take our new knowledge of how air services penetrate existing markets in a country and apply that to potential new routes across all the combinations of airports. That's the key that really lets us find new routes.

Fusing Mobile Data with Traditional Aviation Databases

Beyond that, we don't see the entire population in the data — in the US, for example, we have 80 million anonymous people in the data out of 230 million, only about 20% of the population. That's a big technical challenge: how do we scale up? We could just multiply by five, but that's simplistic and not very credible.

Fortunately, in aviation we're specialists in air mobility, and there are lots of global databases on air demand. We have data from OAG on which aircraft types, with how many seats, are flying on which routes, on which days — down to the hour — going back decades. We have other databases with data on how many people were actually flying on those flights, so we know actual passenger demand. And governments also produce data on air segments, simply by counting people at airports and publishing it.

So we have really high-quality data sources, letting us leverage our existing industry knowledge of air mobility and compare it to what we measure in MobilityMonitor. With a lot of learning and iteration, we figured out how to fuse these datasets — combining aviation knowledge with the new mobility knowledge — to create estimates of total demand at a national level, built up from city-pair observations.

Storytelling and Visualization

How has it been for us? The last part of the process wasn't just the calibration — it was the visualization, and for that I have to say a big thank-you to CARTO. We really needed a powerful tool to share what we found and were discovering with our airline customers, with governments, and with aviation authorities. It's about storytelling — there aren't any existing databases about intercity mobility, so with this analysis we were able to generate new insights, and CARTO's platform has really helped us share them and tell stories about mobility that people aren't used to hearing.

Historically, we would say something like: “We see growth of X million passengers in the regional air segment compared to today's regional air segment — double-digit percentages.” Now, we can say: in India, we see 4.6 billion intercity, all-mode journeys, and only 0.7% of that — 35 million journeys — can be captured for air. It's a complete reorientation of the discussion about how people are moving, and it brings a lot of inspiration to investors, to governments, and to our airlines — who are reassured that they can find demand, for example in Indonesia, where today it isn't always easy for them to find proof that demand exists.

After all, our brand values are all about connecting communities. Having these discussions and talking about mobility and generating community connections has been very valuable for us. That's my story.

Ready to unlock the power of spatial analytics?