Exploring methods for mapping seasonal population changes using mobile phone data
Exploring methods for mapping seasonal population changes using mobile phone data
Data accurately representing the population distribution at the subnational level within countries is critical to policy and decision makers for many applications. Call data records (CDRs) have shown great promise for this, providing much higher temporal and spatial resolutions compared to traditional data sources. For CDRs to be integrated with other data and in order to effectively inform and support policy and decision making, mobile phone user must be distributed from the cell tower level into administrative units. This can be done in different ways and it is often not considered which method produces the best representation of the underlying population distribution. Using anonymised CDRs in Namibia between 2011 and 2013, four distribution methods were assessed at multiple administrative unit levels. Estimates of user density per administrative unit were ranked for each method and compared against the corresponding census-derived population densities, using Kendall’s tau-b rank tests. Seasonal and trend decomposition using Loess (STL) and multivariate clustering was subsequently used to identify patterns of seasonal user variation and investigate how different distribution methods can impact these. Results show that the accuracy of the results of each distribution method is influenced by the considered administrative unit level. While marginal differences between methods are displayed at “coarser” level 1, the use of mobile phone tower ranges provided the most accurate results for Namibia at finer levels 2 and 3. The use of STL is helpful to recognise the impact of the underlying distribution methods on further analysis, with the degree of consensus between methods decreasing as spatial scale increases. Multivariate clustering delivers valuable insights into which units share a similar seasonal user behaviour. The higher the number of prescribed clusters, the more the results obtained using different distribution methods differ. However, two major seasonal patterns were identified across all distribution methods, levels and most cluster numbers: (a) units with a 15% user decrease in August and (b) units with a 20–30% user increase in December. Both patterns are likely to be partially linked to school holidays and people going on vacation and/or visiting relatives and friends. This study highlights the need and importance of investigating CDRs in detail before conducting subsequent analysis like seasonal and trend decomposition. In particular, CDRs need to be investigated both in terms of their area and population coverage, as well as in relation to the appropriate distribution method to use based on the spatial scale of the specific application. The use of inappropriate methods can change observed seasonal patterns and impact the derived conclusions.
Woods, D.
2a542d84-18c1-48d5-b039-ebba67562006
Cunningham, A.
d67452a2-f592-4784-80b2-1bfd8e5f76ae
Utazi, C. E.
e69ca81e-fb23-4bc1-99a5-25c9e0f4d6f9
Bondarenko, M.
1cbea387-2a42-4061-9713-bbfdf4d11226
Lai, Shengjie
b57a5fe8-cfb6-4fa7-b414-a98bb891b001
Rogers, G. E.
68812a4d-98bf-42a5-b21a-593462f34691
Koper, P.
6dc66b8a-6e8e-45c9-9a1b-6322e70a39bf
Ruktanonchai, C. W.
a576fb11-a475-4d48-885a-85938b60a7a8
Erbach-Schoenberg, E. zu
9a1f59b2-c661-42c9-ad94-96772c292add
Tatem, A. J.
6c6de104-a5f9-46e0-bb93-a1a7c980513e
Steele, J.
5cbba8c8-f3fd-41ee-82c8-0aa13c04c04d
Sorichetta, A.
c80d941b-a3f5-4a6d-9a19-e3eeba84443c
28 July 2022
Woods, D.
2a542d84-18c1-48d5-b039-ebba67562006
Cunningham, A.
d67452a2-f592-4784-80b2-1bfd8e5f76ae
Utazi, C. E.
e69ca81e-fb23-4bc1-99a5-25c9e0f4d6f9
Bondarenko, M.
1cbea387-2a42-4061-9713-bbfdf4d11226
Lai, Shengjie
b57a5fe8-cfb6-4fa7-b414-a98bb891b001
Rogers, G. E.
68812a4d-98bf-42a5-b21a-593462f34691
Koper, P.
6dc66b8a-6e8e-45c9-9a1b-6322e70a39bf
Ruktanonchai, C. W.
a576fb11-a475-4d48-885a-85938b60a7a8
Erbach-Schoenberg, E. zu
9a1f59b2-c661-42c9-ad94-96772c292add
Tatem, A. J.
6c6de104-a5f9-46e0-bb93-a1a7c980513e
Steele, J.
5cbba8c8-f3fd-41ee-82c8-0aa13c04c04d
Sorichetta, A.
c80d941b-a3f5-4a6d-9a19-e3eeba84443c
Woods, D., Cunningham, A., Utazi, C. E., Bondarenko, M., Lai, Shengjie, Rogers, G. E., Koper, P., Ruktanonchai, C. W., Erbach-Schoenberg, E. zu, Tatem, A. J., Steele, J. and Sorichetta, A.
(2022)
Exploring methods for mapping seasonal population changes using mobile phone data.
Humanities and Social Sciences Communications, 9 (1), [247].
(doi:10.1057/s41599-022-01256-8).
Abstract
Data accurately representing the population distribution at the subnational level within countries is critical to policy and decision makers for many applications. Call data records (CDRs) have shown great promise for this, providing much higher temporal and spatial resolutions compared to traditional data sources. For CDRs to be integrated with other data and in order to effectively inform and support policy and decision making, mobile phone user must be distributed from the cell tower level into administrative units. This can be done in different ways and it is often not considered which method produces the best representation of the underlying population distribution. Using anonymised CDRs in Namibia between 2011 and 2013, four distribution methods were assessed at multiple administrative unit levels. Estimates of user density per administrative unit were ranked for each method and compared against the corresponding census-derived population densities, using Kendall’s tau-b rank tests. Seasonal and trend decomposition using Loess (STL) and multivariate clustering was subsequently used to identify patterns of seasonal user variation and investigate how different distribution methods can impact these. Results show that the accuracy of the results of each distribution method is influenced by the considered administrative unit level. While marginal differences between methods are displayed at “coarser” level 1, the use of mobile phone tower ranges provided the most accurate results for Namibia at finer levels 2 and 3. The use of STL is helpful to recognise the impact of the underlying distribution methods on further analysis, with the degree of consensus between methods decreasing as spatial scale increases. Multivariate clustering delivers valuable insights into which units share a similar seasonal user behaviour. The higher the number of prescribed clusters, the more the results obtained using different distribution methods differ. However, two major seasonal patterns were identified across all distribution methods, levels and most cluster numbers: (a) units with a 15% user decrease in August and (b) units with a 20–30% user increase in December. Both patterns are likely to be partially linked to school holidays and people going on vacation and/or visiting relatives and friends. This study highlights the need and importance of investigating CDRs in detail before conducting subsequent analysis like seasonal and trend decomposition. In particular, CDRs need to be investigated both in terms of their area and population coverage, as well as in relation to the appropriate distribution method to use based on the spatial scale of the specific application. The use of inappropriate methods can change observed seasonal patterns and impact the derived conclusions.
Text
s41599-022-01256-8
- Version of Record
More information
Submitted date: 25 February 2021
Accepted/In Press date: 29 June 2022
Published date: 28 July 2022
Additional Information:
Funding Information:
The authors would like to thank Mobile Telecommunications Limited for providing access to the mobile phone data. AS, DW, AC, MB, PK, GER, JS are all supported by the Bill & Melinda Gates Foundation (Grant Number OPP1134076). SL is supported by funding from the Bill & Melinda Gates Foundation (OPP1134076, INV-024911), the EU H2020 (MOOD 874850), and the National Natural Science Foundation of China (81773498). AJT is supported by funding from the Bill & Melinda Gates Foundation (OPP1106427, OPP1032350, OPP1134076, OPP1094793), the Clinton Health Access Initiative, the UK Foreign, Commonwealth and Development Office (UK-FCDO), the Wellcome Trust (106866/Z/15/Z, 204613/Z/16/Z), the National Institutes of Health (R01AI160780), and the EU H2020 (MOOD 874850). CEU is supported by funding from the Bill and Melinda Gates Foundation and Gavi, the Vaccine Alliance.
Funding Information:
The authors would like to thank Mobile Telecommunications Limited for providing access to the mobile phone data. AS, DW, AC, MB, PK, GER, JS are all supported by the Bill & Melinda Gates Foundation (Grant Number OPP1134076). SL is supported by funding from the Bill & Melinda Gates Foundation (OPP1134076, INV-024911), the EU H2020 (MOOD 874850), and the National Natural Science Foundation of China (81773498). AJT is supported by funding from the Bill & Melinda Gates Foundation (OPP1106427, OPP1032350, OPP1134076, OPP1094793), the Clinton Health Access Initiative, the UK Foreign, Commonwealth and Development Office (UK-FCDO), the Wellcome Trust (106866/Z/15/Z, 204613/Z/16/Z), the National Institutes of Health (R01AI160780), and the EU H2020 (MOOD 874850). CEU is supported by funding from the Bill and Melinda Gates Foundation and Gavi, the Vaccine Alliance.
Publisher Copyright:
© 2022, The Author(s).
Identifiers
Local EPrints ID: 469270
URI: http://eprints.soton.ac.uk/id/eprint/469270
ISSN: 2662-9992
PURE UUID: 7a223af3-f638-4bf1-b389-38c099160688
Catalogue record
Date deposited: 12 Sep 2022 16:40
Last modified: 06 Jun 2024 02:03
Export record
Altmetrics
Contributors
Author:
G. E. Rogers
Author:
C. W. Ruktanonchai
Author:
E. zu Erbach-Schoenberg
Download statistics
Downloads from ePrints over the past year. Other digital versions may also be available to download e.g. from the publisher's website.
View more statistics