Thursday, September 18, 2014
If Scotland votes ‘Yes’ to independence in the referendum, what will happen to our data?
Sunday, May 6, 2012
Content Analytics - Crossing the Chasm?
That’s where information analytics come in with all their applications such as semantics and auto-classification. Simply put, these technologies are analyzing the actual content of information, deriving insights, meanings, and understanding. It is done by applying powerful algorithms that analyse the content and complete complex tasks such as concept and entity extraction, similarities, trend identification, and sentiment analysis.Therein lies the problem with the technology. The greater the volume of information, the more content analytics are useful - if the volume is small, humans can do it themselves. But to apply analytics on a large volume of information, one needs a significant computing power.
That’s the reason why the actual use of analytics was confined to very few scenarios where money was not an issue. The US intelligence agencies used supercomputers with analytics to weed through millions of intercepted messages from suspected terrorists. Similarly, IBM’s Watson, the Jeopardy winning machine, was a supercomputer. It was fed with encyclopedic knowledge and optimized for a single task: winning Jeopardy. Even though content analytics have been around for well over a decade, most organizations simply could not afford the computing power required.
That may be changing now, thanks to a couple of market shifts. First, Moore’s Law is helping - computing power is becoming more and more affordable and the algorithms are becoming more powerful and efficient.
The other improvement comes from our understanding of what the technology is expected to accomplish. The early requirements for analytics and classification asked for ultra-accuracy. One of the key objections used to be the lack of dependability on automatic classification - if it isn't 100% accurate, it is no good. But today, we understand that the alternative is not perfect by any stretch. The alternative is to rely on humans who are actually pretty pathetic at analyzing and classifying content. In fact, humans are notoriously poor and inconsistent at the job. Getting to 60% accuracy is a typical result of human classification.
That means that a technology that can get us to 80% or even 90% of accuracy is actually much more accurate than any humans and this kind of approach doesn’t require as much computing power as the attempt to reach 99.9% accuracy.
For example, one of the common applications for analytics is legal discovery - the need to quickly produce any electronic evidence requested by a court subpoena. Here the goal today is to produce all the relevant documents and emails with a defensible level of accuracy. Of course we don’t want to pay the expensive lawyers for manual review of thousands of documents. They are paid by the hour - and paid rather well. But we also don’t want to stand accused of failing to produce an important piece of evidence. Until recently, the fear was that unless we can prove that we have electronically discovered all the pertaining documents, the approach would not be defensible in a court of law. And only humans (ehm, lawyers) can guarantee such accuracy - for a hefty fee...
Today, that has changed. The courts increasingly understand the futility of aiming for 100% accuracy and instead accept statistical sampling as a way to confirm accuracy at a reasonable level. After all, both both opposing parties - the plaintiff and the defendant - are in the same boat when it comes down to reviewing a mountain of electronic evidence. That means that the lawyers no longer have to review every document. Instead, statistical evidence of accuracy is considered defensible.
As a result, auto-classification no longer has to aim for 100% accuracy. Instead, a more reasonable level of accuracy backed by statistical sampling has become acceptable - because it is still much better than what humans could ever do manually. And cheaper, of course. That makes analytics much more effective and affordable. With that, analytics are no longer confined to the world of super-computers and finding real-world use cases. Analytics may indeed have crossed the chasm.
Tuesday, June 22, 2010
Man versus Machine
A recent article in Wired Magazine titled “Clive Thompson on the Cyborg Advantage” described the result of a “freestyle” chess tournament in which teams of players competed with help of any computerized aid. What was surprising was that the winner was not the team with a chess grand-master or the team with the most powerful supercomputer. Instead, the team that won was a team that was best able to combine the power of the machine with the human way of thinking.
Years ago, I was dipping into the field of Artificial Intelligence (AI) which was the hype of the time. AI has failed for a variety of reasons. Perhaps it was way ahead of its time but perhaps it attempted to relegate too much decision power to the machines while the human expertise and intuition have always proven superior in the end. And so AI vanished and I’ve moved on to other things like content management.
The problem AI attempted to solve is more than relevant today. Faced with the staggering over-abundance of information, we are trying to find ways in which to use computers to help us make sense of all the data. The first step was making the information retrievable via search. But as soon as we have halfway accomplished that task, we have come to realize that this is not the solution. Virtually every search query produces too many results and the poor humans have to employ their expertise and intuition yet again to weed out the millions of hits.
The next step is to employ machines to automatically analyze and classify the content to reduce the volume of information humans have to deal with. But while such analytics and classification technologies have been around for years, they are still in their infancy. Outside specific applications that deal with limited content volume and scope, we don’t trust the machines yet. Usually, the final decision is up to the humans – just think of the e-Discovery reference model where we find all relevant content and then filter it to reduce the manual review cost. The goal today is to cull the volume of data that humans have to deal with. And that might remain the right approach for some time to come.
The right line of attack might be just like in the freestyle chess match. The solution is to facilitate the best possible interaction between the machines and humans. That needs to be reflected in the software architecture and its user interfaces but perhaps also in the skills required from us, humans. In the near future, it might not be the smartest people who will be most effective but rather those who will be best able to take advantage of the machines to augment their decision making ability.