- Published on
So in my last post I made a claim that taxonomy work was important to resolve vocabulary variations to common concepts, in order to discern patterns from multiple data sources. This article from the New York Times "Computing Crime and Punishment" is a beautiful example of how a thesaurus can be used to recognise similar concepts in unstructured data sources over very long periods of time - in this case, 121 million words describing 197,000 trials over 239 years. Of course vocabularies changed, but Roget's Thesaurus turned out to be a beautiful instrument layered on top of the data.
By Patrick Lambe
- Published on
Really sharp post from Andrew McAfee today on when data analytics trumps expert intuition. Data systems can see patterns over long periods, where human learning from feedback favours short cycles. The critical dependency to my mind however, is that in big data using multiple data sources, you have to have a reliable means to resolve different labels for relevant entities and events to the same concepts. Without that the variability of vocabulary will stymie your attempts to see the patterns over the multiple sources ( and time periods). In case you think this is a minor issue, check out the 2001 deportation of Australian citizen Vivian Alvarez "as an unlawful non-citizen" - all because the various immigration department databases could not resolve different variations of her name.
By Patrick Lambe