Geographical data have always been ‘big’, presenting special challenges for geo-computation. Adding Web-scale ‘Big Data’ makes things even bigger, into the Gigabytes, Petabytes, Terabytes or more…
Speed of processing – even on Adrian’s specially commissioned ‘shedputer’, two 4U IBM System X3850 M2 computers with 24 processor cores, 128GB RAM each and RAID10 solid state disks, pictured in situ, along with a 2U Dell PowerEdge 2950 – was also an issue.
Working with Professor Richard Healey, Gary Burton and David Marshall at the University or Portsmouth, Adrian investigated and assessed the use of High Performance Computers (HPCs) in the Institute of Cosmology and Gravitation’sSCIAMA Supercomputer for Big Data workloads.
Running MapR’s flavour of Hadoop on a 5-node HPC cluster with 48 processor cores, 100GB memory and 6TB disk space and Oracle’s 12c Relational Database Management System (RDBMS) on a 2-node HPC cluster brought significant performance gains for data ingestion, query and analysis.
The old instructions for interfacing MapR (now defunct) with Tableau through Java Database Connectivity (JDBC) connectors is reproduced below, in case it is of any further use to the community.
The supercomputing tests also led to further work with Oracle, including one of the first ‘field tests’ of the Oracle Cloud, including usage of a highly-specified Exadata machine that performed very well with the Oracle 12c RDBMS.
Technology marches on and computing power continues to increase. Adrian’s shedputers still exist but most heavy lifting today is now done in the Cloud…
Adrian’s PhD research had used a number of Natural Language Processing (NLP) systems to search for toponymic (place name) mentions in Online Social Network (OSN) interactions sourced from Twitter and Facebook:
GATE Desktop and GATE Cloud, two long-established open source projects from the University of Sheffield
CLAVIN-rest, a RESTful API built on top of Berico Technologies’ (now Novetta’s) CLAVIN geoparsing and georesolution library.
As part of a separate research project these technologies were applied to Professor Richard Healey’s research into the Illinois Central Rail Road (ICRR) and Chicago and North Western Rail Road (CNWRR).
Digitised County History documents from the Internet Archive were downloaded, e.g.:
Scanned documents were processed with Optical Character Recognition (OCR) software and searched for the words:
railroad
railroads
brakeman
engineer
conductor
fireman
railroad clerk
baggage master
depot master
railroad agent
railroad superintendent
railroad machinist
railroad shops
These terms, relevant to Richard’s research, could be found in the ORC’d texts and used, with some further developments to jump to the relevant pages, to speed up the research process.
Adrian embarked upon Doctoral research in the Department of Geography (as was) at the University of Portsmouth, part-time, in 2011 under the supervision of Professor Richard Healey, who had taught him years earlier as a lecturer on the MSc GIS programme at the University of Edinburgh.
Adrian’s PhD was finally awarded, years later, in 2018. By then the Department of Geography had been renamed the School of the Environment, Geography and Geosciences. That name has since been changed again to the School of the Environment and Life Sciences. And, in the interim, Adrian had decided that his research had much more to do with Computational Social Science than Geography!
Only recently has it become possible, as ‘Web 2.0 desires to read, write, and share personal information’ have developed (Jung, 2015, p53), to know what large numbers of people are saying or thinking as they comment on, share and interact with content online. The rapid growth of Online Social Networks (OSNs) such as Facebook and Twitter has brought billions of users and countless user messages into the public domain. User Generated Content (UGC) now abounds and the traditional communications model of Habermas’ (2011) Public Sphere incorporating governmental, judicial and media power-players appears to have moved towards a more pluralistic model involving overlap between public and newly digitally-enabled ‘private’ spheres (Papacharissi, 2010).
Political opinions expressed online are now widely-made, shared, and increasingly accessible for download from OSN platforms; some records are coordinate-geotagged and many more make frequent toponymic mention of place (Han, Cook, & Baldwin, 2014). For privacy reasons, other than to site operators, IP addresses are not made available in OSN data downloads; removing one of the easiest – although not necessarily accurate (Backstrom, Sun, & Marlow, 2010) – methods to geographically estimate the location of social media communications. If, as Surowiecki (2004) has suggested, the ‘Wisdom of the Crowds’ really can provide more accurate prediction, is it possible that ‘mining’ massive amounts of OSN data for spatiotemporally expressed political opinion could help to detect events, or possibly even ‘call’ an election? After all, as some scholars have suggested, ‘there is a strong relationship between political information [consumption, news seeking] and political participation; the more we know about politics, the more, and more effectively, we participate in political activities’ (Feezell, 2016, p495).
The research was based around:
The opinions, expressed in the public domain, of ~2.4 million users of Online Social Network (OSN) platforms, predominantly of the popular micro-blogging site Twitter, but also of Facebook, have provided the raw material examined in this research. In 1997, 2000 and 2001 Internet users consumed political information from publishers’ websites; today they consume, link to, share and comment on publishers’ material as well as creating original content of their own. The ~8 million records of social media message text, metadata and linked/shared content accessed during the 2012 US Presidential Election and the 2014 Scottish Independence Referendum campaigns have provided an excellent, if at times hard to interpret, politically discursive corpus that has been text and data-mined in various ways.
It proved impossible, unfortunately, to impute localised voting intention from Twitter tweets or Facebook posts but the research did reveal important differences in the communications preferences of different types of users of the two Online Social Networks (OSNs):
Coordinate-geotagging users make fewer toponymic mentions in message text than non-coordinate-geotagging users of two popular OSN platforms;
Coordinate-geotagging users make far fewer URL link shares than non-coordinate-geotagging users, and;
The content of URLs shared by coordinate-geotagging users makes fewer mentions of place than content shared by non-coordinate-geotagging users.
There’s no real substitute for reading the thesis, but at 120,171 words (69,014 excluding references and appendices), it is quite a heavy read.
For a condensed summary, based on a presentation of the research at the PLATIAL ’19 conference, held at the University of Warwick on 5th September 2019, please flick through the presentation below.
The original slides, in Microsoft PowerPoint format, may also be downloaded.
The completion of his PhD, something Adrian had returned to later in life, after an 18 year career outside academia in business, has been a notable achievement.
For anyone else considering Doctoral research the challenge is recommended. It is probable, however, that starting younger and working on things full time makes for a somewhat easier approach…