Showing posts with label classification. Show all posts
Showing posts with label classification. Show all posts

Thursday, September 18, 2014

If Scotland votes ‘Yes’ to independence in the referendum, what will happen to our data?

The following article originally appeared in the IT Pro Portal on September 18, 2014:

Scotland goes to the polls today to vote on independence from the Union. If a ‘Yes’ vote is passed, it will throw into question the massive issue of data sovereignty.

It’s a curious notion, and one that both Whitehall and Holyrood have not publicly answered. If we consider the consequences from a data protection perspective, they are incredibly complex. Especially as the EU Data Protection Directive mandates that data cannot be transferred outside of the 28-member states territory. This means that organisations need to prove where their data is at all times. Scotland, as part of the UK, is presently an EU member but this could quickly change, as only last night the Telegraph reported that Spain’s European Foreign Affairs Minister said that a separate Scotland would need to wait five years for EU membership and join the single currency.

The next wave of EU data protection reforms will introduce further enforcements around information crossing borders. The fact that the majority of business communications are digital by nature – such as email and productivity tools like Microsoft Word and Excel – and are effectively borderless, these carry more problems for data sovereignty compliance. When these reforms were first suggested last year, the Direct Marketing Association said it was ‘”strict and unworkable” and claimed that it would cost UK businesses an eye-watering £47 billion in lost sales and regulatory costs.

Today your data may be based in the UK, but tomorrow, will it be in England or Scotland? The key is allowing companies from either country to pick where they want their data to live and guarantee that it resides there.

A ‘Yes’ vote will mean that the data will have to be migrated. But to where at this stage can only be speculated. What we do know is that for companies that want to move their data from Scotland to England, capacity and power availability within the Greater London M25 belt will be under extreme pressure. If Scottish data needs to head north, then perhaps the challenge isn’t quite so unwieldy. Scotland’s climate is well-suited for cooling power-hungry server farms. It boasts a thriving data centre economy, with substantial investment pumped into the hundreds across the country, with some areas even building new locally-generated renewable energy-based data centres, which are expected to go online next year. Scotland’s digital sector is worth around £3 billion to the economy and boasts over 73,000 jobs. For a population of just over five million, it’s certainly a healthy industry.

The nationalist campaigners suggest they could transfer Scotland’s data from Whitehall systems by 2018, but this is likely to result in considerable disruption to public services, not to mention commercial implications for organisations that own or host from data centres based there. Over time, these issues will have to be remedied. The result will be that data sovereignty will become a board issue and part of future business and legal operations as principle. But this is not necessarily a bad thing.

The recent spotlight on data sovereignty originates from the much-reported WikiLeaks and Snowden affairs and the US’ National Security Agency spying revelations. These stories have created such a wake that, according to ResearchNow, a quarter of UK companies are now expected to pull their data out of US data centres. Protecting the integrity of data is definitely at the top of the corporate agenda and it requires sovereignty and security embedded by design.

If Scotland does separate form the UK and takes a while to decide how it wants to pursue its membership with the EU, it will mean that all data housed in Scotland from another country’s origin would need to move inside the EU. An alternative is that a provision could be made for its own data protection regulation but this would need to be written and ratified – a pretty costly and complicated exercise, never mind the process of data migration.

Given that both England and Scotland speak the same language means that the information doesn’t need to be translated, but that also makes it more difficult to separate Scottish data from English data. The challenge will be organising and sorting through Petabytes of data and establishing whether it originates from England or Scotland. There are obvious clues, such as place names and cultural references, that will help with labelling but in reality this particular job is for humans who unfortunately are inherently unreliable when it comes to organising information. Auto-classification tools on the other hand typically deliver 80–90 per cent accuracy, as opposed to human classification, that on average results in 60 per cent of information being properly classified.

For the whole of the United Kingdom to adhere to data protection laws and deal with the sovereignty of data requires a strategic view when it comes to managing enterprise information. As far as Scottish, Welsh and English borders are concerned, the task of migrating so much information will be a tremendous undertaking. Let’s hope we don’t get to that stage.

Sunday, May 6, 2012

Content Analytics - Crossing the Chasm?

The overabundance of information is one of the greatest challenges today. Things have sure changed a bit since the days of the 90s when Microsoft used to promise us a PC on every desk and information at our fingertips. Today, we have plenty of information at our fingertips. In fact, we have so much information coming at us from all sides that making sense of it became one of the key information management challenges.

That’s where information analytics come in with all their applications such as semantics and auto-classification. Simply put, these technologies are analyzing the actual content of information, deriving insights, meanings, and understanding. It is done by applying powerful algorithms that analyse the content and complete complex tasks such as concept and entity extraction, similarities, trend identification, and sentiment analysis.

Therein lies the problem with the technology. The greater the volume of information, the more content analytics are useful - if the volume is small, humans can do it themselves. But to apply analytics on a large volume of information, one needs a significant computing power.

That’s the reason why the actual use of analytics was confined to very few scenarios where money was not an issue. The US intelligence agencies used supercomputers with analytics to weed through millions of intercepted messages from suspected terrorists. Similarly, IBM’s Watson, the Jeopardy winning machine, was a supercomputer. It was fed with encyclopedic knowledge and optimized for a single task: winning Jeopardy. Even though content analytics have been around for well over a decade, most organizations simply could not afford the computing power required.

That may be changing now, thanks to a couple of market shifts. First, Moore’s Law is helping - computing power is becoming more and more affordable and the algorithms are becoming more powerful and efficient.

The other improvement comes from our understanding of what the technology is expected to accomplish. The early requirements for analytics and classification asked for ultra-accuracy. One of the key objections used to be the lack of dependability on automatic classification - if it isn't 100% accurate, it is no good. But today, we understand that the alternative is not perfect by any stretch. The alternative is to rely on humans who are actually pretty pathetic at analyzing and classifying content. In fact, humans are notoriously poor and inconsistent at the job. Getting to 60% accuracy is a typical result of human classification.

That means that a technology that can get us to 80% or even 90% of accuracy is actually much more accurate than any humans and this kind of approach doesn’t require as much computing power as the attempt to reach 99.9% accuracy.

For example, one of the common applications for analytics is legal discovery - the need to quickly produce any electronic evidence requested by a court subpoena. Here the goal today is to produce all the relevant documents and emails with a defensible level of accuracy. Of course we don’t want to pay the expensive lawyers for manual review of thousands of documents. They are paid by the hour - and paid rather well. But we also don’t want to stand accused of failing to produce an important piece of evidence. Until recently, the fear was that unless we can prove that we have electronically discovered all the pertaining documents, the approach would not be defensible in a court of law. And only humans (ehm, lawyers) can guarantee such accuracy - for a hefty fee...

Today, that has changed. The courts increasingly understand the futility of aiming for 100% accuracy and instead accept statistical sampling as a way to confirm accuracy at a reasonable level. After all, both both opposing parties - the plaintiff and the defendant - are in the same boat when it comes down to reviewing a mountain of electronic evidence. That means that the lawyers no longer have to review every document. Instead, statistical evidence of accuracy is considered defensible.

As a result, auto-classification no longer has to aim for 100% accuracy. Instead, a more reasonable level of accuracy backed by statistical sampling has become acceptable - because it is still much better than what humans could ever do manually. And cheaper, of course. That makes analytics much more effective and affordable. With that, analytics are no longer confined to the world of super-computers and finding real-world use cases. Analytics may indeed have crossed the chasm.

Tuesday, June 22, 2010

Man versus Machine

A recent article in Wired Magazine titled “Clive Thompson on the Cyborg Advantage” described the result of a “freestyle” chess tournament in which teams of players competed with help of any computerized aid. What was surprising was that the winner was not the team with a chess grand-master or the team with the most powerful supercomputer. Instead, the team that won was a team that was best able to combine the power of the machine with the human way of thinking.

Years ago, I was dipping into the field of Artificial Intelligence (AI) which was the hype of the time. AI has failed for a variety of reasons. Perhaps it was way ahead of its time but perhaps it attempted to relegate too much decision power to the machines while the human expertise and intuition have always proven superior in the end. And so AI vanished and I’ve moved on to other things like content management.

The problem AI attempted to solve is more than relevant today. Faced with the staggering over-abundance of information, we are trying to find ways in which to use computers to help us make sense of all the data. The first step was making the information retrievable via search. But as soon as we have halfway accomplished that task, we have come to realize that this is not the solution. Virtually every search query produces too many results and the poor humans have to employ their expertise and intuition yet again to weed out the millions of hits.

The next step is to employ machines to automatically analyze and classify the content to reduce the volume of information humans have to deal with. But while such analytics and classification technologies have been around for years, they are still in their infancy. Outside specific applications that deal with limited content volume and scope, we don’t trust the machines yet. Usually, the final decision is up to the humans – just think of the e-Discovery reference model where we find all relevant content and then filter it to reduce the manual review cost. The goal today is to cull the volume of data that humans have to deal with. And that might remain the right approach for some time to come.

The right line of attack might be just like in the freestyle chess match. The solution is to facilitate the best possible interaction between the machines and humans. That needs to be reflected in the software architecture and its user interfaces but perhaps also in the skills required from us, humans. In the near future, it might not be the smartest people who will be most effective but rather those who will be best able to take advantage of the machines to augment their decision making ability.