menu

Showing posts with label Bayes Knowledge. Show all posts
Showing posts with label Bayes Knowledge. Show all posts

Wednesday, 14 March 2018

Wednesday, 1 June 2016

Bayesian networks for Cost, Benefit and Risk Analysis of Agricultural Development Projects


Successful implementation of major projects requires careful management of uncertainty and risk. Yet, uncertainty is rarely effectively calculated when analysing project costs and benefits. In the case of major agricultural and other development projects in Africa this challenge is especially important.

A paper just published* in the journal Experts Systems with Applications presents a Bayesian network (BN) modelling framework to calculate the costs, benefits, and return on investment of a project over a specified time period, allowing for changing circumstances and trade-offs. Marianne Gadeberg and Eike Luedeling have written an overview of the work here.

The framework uses hybrid and dynamic BNs containing both discrete and continuous variables over multiple time stages. The BN framework calculates costs and benefits based on multiple causal factors including the effects of individual risk factors, budget deficits, and time value discounting, taking account of the parameter uncertainty of all continuous variables. The framework can serve as the basis for various project management assessments and is illustrated using a case study of an agricultural development project. The work was a collaboration between the World Agroforestry Centre (ICRAF), Nairobi, Kenya, the Risk Information Management Group at Queen Mary (as part of the BAYES-KNOWLEDGE project) and Agena Ltd.

*The full reference is:
Yet, B., Constantinou, A., Fenton, N., Neil, M., Luedeling, E., & Shepherd, K. (2016). "A Bayesian Network Framework for Project Cost, Benefit and Risk Analysis with an Agricultural Development Case Study" . Expert Systems with Applications, Volume 60, 30 October 2016, Pages 141–155. DOI: 10.1016/j.eswa.2016.05.005
Until July 2016 the full published pdf is available for free.  A permanent pre-publication pdf is available here.

See also: Can we build a better project: assessing complexities in development projects

Acknowledgements: Part of this work was performed under the auspices of EU project ERC-2013-AdG339182-BAYES_KNOWLEDGE and part under ICRAF Contract No SD4/2012/214 issued to Agena. We acknowledge support from the Water, Land and Ecosystems (WLE) program of the Consultative Group on International Agricultural Research (CGIAR).

Thursday, 26 May 2016

Using Bayesian networks to assess new forensic evidence in an appeal case


If new forensic evidence becomes available after a conviction how do lawyers determine whether it raises sufficient questions about the verdict in order to launch an appeal? It turns out that there is no systematic framework to help lawyers do this. But a paper published today by Nadine Smit and colleagues in Crime Science presents such a framework driven by a recent case, in which a defendant was convicted primarily on the basis of sound evidence, but where subsequent analysis of the evidence revealed additional sounds that were not considered during the trial.

From the case documentation, we know the following:
  • A baby was injured during an incident on the top floor of a house
  • Blood from the baby was found on the wall in one of the rooms upstairs
  • On an audio recording of the emergency telephone call made by the suspect, a scraping sound (allegedly indicating scraping blood off a wall) can be heard
  • The suspect was charged with attempted murder 
The audio evidence played a significant role in the trial. But, during the appeal preparation process, the call was re-analysed by an audio expert on behalf of the defence, and four other sounds were identified on the same recording that, according to the expert, showed similarities to the original sound. In particular, one of these sounds was of interest because of background noise that could be heard simultaneously. The background noise was presumed to be the television, which was located in a different room to where the prosecution argued the scraping of the blood took place.  During this second sound, the TV (located downstairs) could be heard simultaneously on the emergency recording. A statement by the police reads that the suspect was frequently rubbing his face in their presence. The defence proposed that the incriminating sound in the recording was not blood scraping after all, but simply the defendant rubbing his face.

The framework described in Smit's paper is intended to overcome the gap between what is generally known from scientific analyses and what is hypothesized in a legal setting. It is based on Bayesian networks (BNs) which are a structured and understandable way to evaluate the evidence in the specific case context and present it in a clear manner in court. However, BN methods are often criticised for not being sufficiently transparent for legal professionals. To address this concern the paper shows the extent to which the reasoning and decisions of the particular case can be made explicit and transparent. The BN approach enables us to clearly define the relevant propositions and evidence, and uses sensitivity analysis to assess the impact of the evidence under different prior assumptions. The results show that such a framework is suitable to identify information that is currently missing, and clearly crucial for a valid and complete reasoning process. Furthermore, a method is provided whereby BNs can serve as a guide to not only reason with incomplete evidence in forensic cases, but also identify very specific research questions that should be addressed to extend the evidence base to solve similar issues in the future.

Full citation:
Smit, N. M., Lagnado, D. A., Morgan, R. M., & Fenton, N. E. (2016). "An investigation of the application of Bayesian networks to case assessment in an appeal case". Crime Science, 2016, 5: 9, DOI 10.1186/s40163-016-0057-6 (open source). Published version pdf.
The research was funded by the Engineering and Physical Sciences Research Council of the UK through the Security Science Doctoral Research Training Centre (UCL SECReT) based at University College London (EP/G037264/1), and the European Research Council (ERC-2013-AdG339182-BAYES_KNOWLEDGE). 

The BN model (which is fully spceified in the paper) was built and run using the free version of AgenaRisk.

Tuesday, 8 March 2016

Anthony Constantinou presents BAYES-KNOWLEDGE work in Parliament


Anthony Constantinou and his poster

Updated 8 March 2016

Anthony Constantinou was selected to present his work (summarised in this poster) on the BAYES-KNOWLEDGE project in Parliament on 7 March 2016 as part of the SET for Britain awards 2016 This was the Press Release from Parliament:

Dr Anthony Constantinou, 31, a Post-Doctoral Researcher at Queen Mary University of London, hailing from Limassol, Cyprus, is attending Parliament to present his mathematics research to a range of politicians and a panel of expert judges, as part of SET for Britain on Monday 7 March.

Anthony’s poster on research about Decision Systems which are based on probabilistic graphical models for uncertainty quantification and risk management, will be judged against dozens of other scientists’ research in the only national competition of its kind. Anthony is funded as part of the BAYES-KNOWLEDGE project (http://bayes-knowledge.com/).

Anthony was shortlisted from hundreds of applicants to appear in Parliament.

On presenting his research in Parliament, he said, “This is a great opportunity to demonstrate the art of intelligent decision making, which is merely based on identifying the best possible set of actions to be taken on the basis of trade-offs in an effort to achieve some objectives; then having to explain why an objective has not been met! Having applied these methods to a diverse range of real-world application domains, spanning from forensics and medical sciences to sports and gambling market efficiency, I do hope I can keep the audience entertained.”

Stephen Metcalfe MP, Chairman of the Parliamentary and Scientific Committee, said:

“This annual competition is an important date in the parliamentary calendar because it gives MPs an opportunity to speak to a wide range of the country’s best young researchers.

“These early career engineers, mathematicians and scientists are the architects of our future and SET for Britain is politicians’ best opportunity to meet them and understand their work.”

Anthony’s research has been entered into the Mathematical Sciences session of the competition, which will end in a gold, silver and bronze prize-giving ceremony.

Judged by leading academics, the gold medalist receives £3,000, while silver and bronze receive £2,000 and £1,000 respectively.

The Parliamentary and Scientific Committee runs the event in collaboration with the Royal Academy of Engineering, the Royal Society of Chemistry, the Institute of Physics, the Royal Society of Biology, The Physiological Society and the Council for the Mathematical Sciences, with financial support from Essar, the Clay Mathematics Institute, Warwick Manufacturing Group (WMG), the Institute of Biomedical Science, the Bank of England and the Society of Chemical Industry.
Although Anthony did not win one of the final prizes his entry was highly praised by the judges and many MPs who he spoke with.

Anthony about to enter the Houses of Parliament
Anthony with Nick Clegg and David Cameron!
 
Anthony with poster

The prize giving bit with Science Minister Joe Johnson (5th from right)


Improving Bayesian networks by learning from similar ones


A major challenge of the BAYES-KNOWLEDGE project is about how to build useful and accurate Bayesian network (BN) models for decision support when there is little relevant data. Much of what we have done involves exploiting expert judgment. But my colleagues Yun Zhou and Tim Hospedales have developed a so-called 'tranfer learning' method that enables us to leverage data from different but related problems. Suppose, for example, we have a BN model for a particular medical diagnostic problem that we built based on limited data and expert judgment in the UK. But suppose also that a model for the same (or very similar) diagnostic problem has been developed in the USA based on a much larger data set. Some assumptions in the US model will be different to the UK (such as the population demographics or particular testing methods) but much of the underlying pathology will be the same. The challenge is to understand and exploit the heterogeneous relatedness of the models.

The result of this work has just been published in an article in the Elsevier journal Expert Systems with Applications. Elsevier have provided free access to the article until April 24, 2016:

http://authors.elsevier.com/a/1SfM-3PiGT01St
 
The article describes a new transfer learning algorithm for improved BN parameter learning, and the  experimental results demonstrate its superiority compared to other state-of-the-art parameter transfer methods. The method is applied to a real-world medical case study, namely the problem of trauma care (a problem for which our team had initially developed a decision support BN model in collaboration with UK medics).

Full reference for the new article:

Zhou, Y., Hospedales, T., Fenton, N. E. (2016), "When and where to transfer for Bayes net parameter learning", Expert Systems with Applications. 55, 361-373
http://dx.doi.org/10.1016/j.eswa.2016.02.011.

Sunday, 18 October 2015

What is the value of missing information when assessing decisions that involve actions for intervention?

This is a summary of the following new paper:

Constantinou AC, Yet B, Fenton N, Neil M, Marsh W  "Value of Information analysis for interventional and counterfactual Bayesian networks in forensic medical sciences". Artif Intell Med. 2015 Sep 8  doi:10.1016/j.artmed.2015.09.002. The full pre-publication version can be found here.

Most decision support models in the medical domain provide a prediction about a single key unknown variable, such as whether a patient exhibiting certain symptoms is likely to have (or develop) a particular disease.

However we seek to enhance decision analysis by determining whether a decision based on such a prediction could be subject to amendments on the basis of some incomplete information within the model, and whether it would be worthwhile for the decision maker to seek further information prior to the decision. In particular we wish to incorporate interventional actions and counterfactual analysis, where:
  • An interventional action is one that can be performed to manipulate the effect of some desirable future outcome. In medical decision analysis, an intervention is typically represented by some treatment, which can affect a patient’s health outcome.
  • Counterfactual analysis enables decision makers to compare the observed results in the real world to those of a hypothetical world; what actually happened and what would have happened under some different scenario.
The method we use is based on the underlying principle of Value of Information. This is a technique initially proposed in economics for the purposes of determining the amount a decision maker would be willing to pay for further information that is currently unknown within the model.





The type of predictive decision support models to which our work applies are Bayesian networks. These are graphical models which represent the causal or influential relationships between a set of variables and which provide probabilities for each unknown variable.

The method is applied to two real-world Bayesian network models that were previously developed for decision support in forensic medical sciences. In these models a decision maker (such as a probation officer or a clinician) has to determine whether to release a prisoner/patient based on the probability of the (unknown) hypothesis variable: “individual violently reoffends after release”. Prior to deciding on release, the decision maker has the option to simulate various interventions to determine whether an individual’s risk of violence can be managed to acceptable levels. Additionally, the decision maker may have the option to gather further information about the individual. It is possible that knowing one or more of these unobserved factors may lead to a different decision about release.

We used the method to examine the average information gain; that is, what we learn about the importance of the factors that remain unknown within the model. Based on six different sets of experiments with various assumptions we show that:
  1. the average relative percentage gain in terms of Value of Information ranged between 11.45% and 59.91% (where a gain of X% indicates an expected X% relative reduction of the risk of violent reoffence);
  1. the potential amendments in Decision Making, as a result of the expected information gain, ranged from 0% to 86.8% (where an amendment of X% indicates that X% of the initial decisions are expected to have been altered).
The key concept of the method is that if we had known that the individual was, for example, a substance misuser, we would have arranged for a suitable treatment; whereas without having information about substance misuse it is impossible to arrange such a treatment and, thus, we risk not treating the individual in the case where he or she is a substance misuser.

The method becomes useful for decision makers, not only when decision making is subject to amendments on the basis of some unknown risk factors, but also when it is not. Knowing that a decision outcome is independent of one or more unknown risk factors saves us from seeking information about that particular set of risk factors.

This summary can also be found on the Atlas of Science

Thursday, 15 October 2015

Talk: Bayesian networks: why smart data is better than big data

by Prof. Norman Fenton from the School of Electronic Engineering and Computer Science (QMUL)
WHEN: Fri, 16th October 2 - 3 pm
WHERE: People's Palace PP2 (Mile End Campus)

"This talk will provide an introduction to Bayesian networks which, due to relatively recent algorithmic breakthroughs, has become an increasingly popular technique for risk assessment and decision analysis. I will provide an overview of successful applications (including transport safety, medical, law/forensics, operational risk, and football prediction). What is common to all of these applications is that the Bayesian network models are built using a combination of expert judgment and (often very limited) data. I will explain why Bayesian networks ‘learnt’ purely from data – even when ‘big data’ is available - generally do not work well."

All are welcome. The seminar consists of an app. 45 min long lecture and discussion.
In case of any questions, feel free to contact me.
Hope to see you tomorrow,

Judit Petervari
____________________
Judit Petervari
PhD Student

Biological and Experimental Psychology Group
School of Biological and Chemical Sciences
Queen Mary University of London
Mile End Road
E1 4NS London
United Kingdom

E-mail: j.petervari@qmul.ac.uk
Office: G.E. Fogg Building, Room 2.16

Tuesday, 15 September 2015

Yet another flawed statistical study attracts massive unquestioning attention


The Guardian, 29 Sept 2015
A very widely reported story in today’s news (see, for example, the report in the Guardian and this Press release) claims that companies in which there is at least one female executive on the Board (‘gender diverse’ companies) in the US, UK and India outperform companies with male-only executives by a staggering US$655 billion per year. The story is based on a study by Grant Thornton whose representative Francesca Lagerberg concludes:
“The research clearly shows what we have been talking about for a while: that diversity leads to better decision-making”.
As is typical when the results of a statistical study fit a popular narrative, the story attracted massive, unquestioning attention. Unfortunately, while I am sure that most people agree that greater gender diversity in the Boardroom is a worthy objective, based on the ‘full report’ – and in the absence of other data - Lagerberg's claim is simply not supported. In fact, the study exemplifies some of the classic misuses of statistics that we wrote about in the first chapter of our book and highlights yet again the need for proper causal/explanatory models to be used in statistical studies such as these*.

Moreover, using the data in Lagerberg's study it is possible to construct a simple causal model (a Bayesian network) that replicates the results but with provably opposite conclusions: diversity decreases performance.

The full report and BN model are provided here. The model can be run in the free version of AgenaRisk.

*Making such an approach both universally feasible and acceptable is the major objective of BAYES-KNOWLEDGE.

Sunday, 30 August 2015

Using Bayesian networks to assess and manage risk of violent reoffending among prisoners

Fragment of BN model
Probation officers, clinicians, and forensic medical practitioners have for several years sought improved decision support for determining whether and when to release prisoners with mental health problems and a history of violence.  It is critical that the risk of violent re-offending is accurately measured and, more importantly, well managed with causal interventions to reduce this risk after release. The well-established 'risk predictors' in this area of research are typically based on statistical regression models and their results are less than convincing. But recent work undertaken at Queen Mary University of London has resulted in Bayesian network (BN) models that not only have much greater accuracy, but which are also much more useful for decision support. The work has been developed as part of a collaboration between the Risk and Information Management group and the medical practitioners of the Violence Prevention Research Unit (VPRU) of the Wolfson Institute of Preventative Medicine.

The (BN) model, called DSVM-P (Decision Support for Violence Management – Prisoners) captures the causal relationships between risk factors, interventions and violence.  It also allows for specific risk factors to be targeted for causal intervention for risk management of future re-offending. These decision support features are not available in the previous generation of models used by practitioners and forensic psychiatrists.

Full reference:
Constantinou, A., Freestone M., Marsh, W., Fenton, N. E. , Coid, J. (2015) "Risk assessment and risk management of violent reoffending among prisoners", Expert Systems With Applications 42 (21), 7511-7529.  Published version: http://dx.doi.org/10.1016/j.eswa.2015.05.025.
Download Pre-publication draft.

Friday, 17 July 2015

The use of Bayes in the Netherlands Appeal Court


Henry Prakken
Norman Fenton, 17 July 2016

There has been an important development on the use of Bayes in the Law in the Netherlands, with what is possibly the first full Bayesian analysis of a major crime in an appeal court there.

The case, referred to as the “Breda 6”, was the 1993 murder of a Chinese woman in Breda in her son’s restaurant. Six young people were convicted of the crime and sentenced to up to 10 years in jail (all have since completed their sentences).  In 2012 the advocate general recommended the case be looked at again because it centred on confessions which may have been false.

In the review of the case a Bayesian argument supporting the prosecution case was presented by Frans Alkemade (Update: see below about some concerns Frans has about this article). Frans is actually a physicist who previously used Bayes to analyse a drugs trafficking case, concluding in a report commissioned by the prosecution, that there was not enough evidence for a conviction (the suspect was acquitted). The court requested that Henry Prakken (professor in Legal Informatics and Legal Argumentation at the University of Groningen) respond to the Bayesian argument.

In June, while he was preparing his response, I met Henry at the International Conference on AI and the Law in San Diego. Henry told me about the case and Alkemade's analysis for which the guilty hypothesis was "At least some of the six suspects were involved in the crime, which took place after 4:30 on the night of 3 and 4 july 1993, and which included the luring of the victim to the restaurant by at least some of the female suspects". Alkemade interpreted "involved" in a weak way and Henry said:
"..it could even be no more than just knowing about the crime. In fact, one of my points of criticism was that this guilt hypothesis is not useful for the court, since it is consistent with the innocence of any of the individual suspects (and even with the collective innocence of all three male suspects)."
Among other things, Alkemade focused on two pieces of evidence:
  1. A report by the Criminal Intelligence Unit (CID) of the Dutch police, saying that they had received information from "usually reliable" sources identifying two of the male defendants and one of the female defendants as being involved in the murder**.
  2. The subsequent discovery that two of the female defendants (not mentioned in the CID report and who supposedly knew the three defendants mentioned in the CID report) worked next door to the murder scene.  
From Henry's description of the analysis, it seemed that Alkemade did not account for all relevant unknown variables and dependencies*** (also see the update), such as the possibility that the anonymous tip-off may have been both malicious and prompted by the fact that the caller knew the defendants worked next door to the murder scene (making the tip-off more believable). This would mean that the combination of the two pieces of evidence would not have been such an incredible coincidence if the defendants were innocent. So in that sense the Bayesian argument was over-simplistic. On the other hand it was also too complex for lawyers to understand since it was presented 'from first principles' in the sense that all of the detailed Bayesian inference calculations were spelled out. For the reasons we have explained in detail here it seemed like a Bayesian network (BN) model would be far more suitable. I therefore produced - in discussions with Henry - a generic BN model to reason about the impact of anonymous evidence when combined with other evidence that can influence the anonymous tip-off (the model is here and can be run using the free AgenaRisk software).

The intention was not to replicate all of the features of the case but rather to demonstrate the impact of missing dependencies in Alkemade's argument. Indeed, with a range of reasonable assumptions, the BN model pointed to a much lower probability of guilt than suggested by Alkemade's calculations.

Henry presented his response in court last week. He said:
"The court session was sometimes frustrating, since the discussion was fragmentary and sometimes I had the impression that the court felt it was confronted with a battle of the experts without the means to understand who was right."
In my view this case confirms our claim that presenting a Bayesian legal argument from first principles (as Alkemade did) is not a good idea. The very fact that people assume it is necessary to do this for Bayes to be accepted is actually the reason (ironically) that there will continue to be very strong resistance to accepting Bayes in the courtroom. Why? Because it means you are restricted to ludicrously over-simplistic (and normally flawed) models of the case (3 unknowns maximum) because that is the limit of the Bayesian calculations you can do by hand and explain clearly. Our proposed solution is to model (and run) the case properly using a BN and report back on the results stating in lay terms what the model assumptions were and how sensitive the conclusions are to different prior assumptions.

**Henry told me that there was something funny with the CID report, as also noted by the advocate general. According to the CID report, the anonymous informant had also accused the defendants mentioned in the report of committing several other crimes, but in the investigations preceding the revision case the police investigators had not been able to find any confirmation of these other crimes, not even reports of these supposed crimes to the police by the supposed victims. As the advocate general stated in 2012, this casts doubt on the reliability of the CID informants.

***My colleague Richard Gill - who knows Alkemade - says that Alkemade was careful to define his pieces of evidence in such a way that he thinks that he can justify the independence assumptions which he needs in order to at least conservatively bound the Likelihood ratio coming from each piece of evidence in turn.

UPDATE 21 July 2015: Frans has contacted me stating a number of concerns about the above narrative and provided a number of technical insights that I was not aware of. As an expert witness in a case that is still under trial, he does not feel free to discuss any details in public, but once the trial has finished I will provide an updated report that incorporates his comments. What I can confirm, however, is that in order to do the calculations manually Frans could not model dependencies between different pieces of evidence - a major limitation - although he did make clear the limitations and pitfalls of this in his report.

See also:

Thursday, 9 July 2015

Why target setting leads to poor decision-making


Norman Fenton is the co-author of an article in Nature published today that addresses the issue of improved decision-making in the context of international sustainable development goals. The article pushes for a Bayesian, smart-data approach:

We contend that target-setting is flawed, costly and could have little — or even negative — impact. First, targets may have unintended consequences. For example, education quality as a whole suffered in some countries that diverted resources to early schooling to meet the target of the Millennium Development Goal (MDG) of achieving universal primary education.

Second, target-setting inhibits learning by focusing efforts on meeting the target rather than solving the problem. The milestones are easily manipulated — aims such as halving deaths from road-traffic accidents can trigger misreporting if the performance falls short or encourage underperformance if the goal can be exceeded.

Third, it is costly: development partners will have to reallocate scant resources for a 'data revolution' that will cost an estimated US$1 billion a year.

We advocate a different approach. Governments and the development community need to embrace decision-analysis concepts and tools that have been used for decades in mining, oil, cybersecurity, insurance, environmental policy and drug development.
The approach is based on five principles:
  1. Replace targets with measures of investment return
  2. Model intervention decisions
  3. Integrate expert knowledge
  4. Include uncertainty in predictive models
  5. Measure the most informative variables
Recommendations include the following:
It is a common mistake to assume that 'evidence' is the same as 'data' or that 'subjective' means 'uninformative'. Decision-making should draw on all appropriate sources of evidence. In developing countries where data are sparse, expert knowledge can fill the gaps. For instance, in our assessment of the viability of agroforestry projects in Africa, we used our experience to set ranges on tree-survival rates, costs of raising tree seedlings and farm prices of tree products.
 ....
Decision theorists and local experts will have to work together to identify relevant variables, causal associations and uncertainties. The most widely accepted method of incorporating knowledge for probability assessment is Bayes' theorem. This updates the likelihood of a belief in some event (such as whether an intervention will reduce poverty) when observing new evidence about the event (such as the occurrence of drought). Bayesian analyses — incorporating historical data and expert judgement — are used in transport and systems-safety assessments, medical diagnosis, operational risk assessment in finance and in forensics, but seldom in development. They should be used, for example, to evaluate the relative risks of competing development interventions. 
 ....
Decision-makers .. should employ probabilistic decision analysis, for example Monte Carlo simulations or Bayesian network models. Provided that such models are developed using properly calibrated expert judgement and decision-focused data, they can incorporate the key factors and outcomes and the causal relationships between them. For instance, simulations for evaluating options for building a water pipeline could take into account rare 'what-if' scenarios, such as a hurricane during development, and predict (with probabilities) the time and cost of implementation and the benefits of improved water supply.

Thursday, 26 March 2015

The risk of flying


Norman Fenton, 26 March 2015

I have just done an interview on BBC Radio Scotland about aircraft safety in the light of the GermanWings crash - which now appears to have been a deliberate act of sabotage by the co-pilot*. I have uploaded a (not very good) recording of it here (mp3 file - it is just under 4 minutes) or here (a more compact m4a file)

Because this type of event is so rare classical frequentist statistics provides no real help when it comes to risk assessment. In fact, it is exactly the kind of risk assessment problem for which you need causal models and expert judgement (as explained in our book) if you want any kind of risk insights.

Irrespective of this particular incident, the interview gave me the opportunity to highlight a very common myth, namely that “flying is the safest form of travel”.  If you look at deaths per million travellers then, indeed, there are 50 times as many car deaths as plane deaths. However, this is a silly measure because there are so many more car travellers than plane travellers. So, typically, analysts use deaths per million miles travelled; with respect to this measure car travel is still 'riskier' than air travel, but the death rate is only about twice as high as plane deaths. But this measure is also biased in favour of planes because the average plane journey is much further than the average car journey.

So a much fairer measure is the number of deaths per passenger journey. And for this, the rate of plane deaths is actually three times higher than car deaths; in fact only bikes and motorbikes are worse than planes.

Despite all this there is still a very low probability of a plane journey resulting in fatalities - about 1 in half a million (and much less on commercial flights in Western Europe). However, if we have reason to believe that, say, recent converts to a terrorist ideology have been training and becoming pilots then the probability of the next plane journey resulting in fatalities becomes much higher, despite the past data.

*I had an hour’s notice of the interview and was told what I would be asked.  I was actually not expecting to be asked about how to assess the risk of this specific type of incident;  I was assuming I would only be asked about aircraft safety risk in general and about the safety record of the A320.

Postscript: Following the interview a colleage asked:
"Did you have the mental issues of the co-pilot on the radar when you replied? "
My response: Interesting question. A few years back we were involved extensively in work with NATS (National Air Traffic Safety) to model/predict risk of mid-air collision over the UK airspace. In particular NATS wanted to know how the probability of a mid-air collision might change given different proposals for changes to the ATM architecture (e.g. ‘adding new ground radar stations’ versus ‘adding new on-board collisions alert systems’). Now - apart from three incidents in the late 1940’s which all involved at least one military jet - there has not been any actual mid-air collisions over UK airspace (so negligible data there) and the proposed technology was ‘new’ (so no directly relevant data there) but there was a LOT of data on "near misses" of different degrees of seriousness  and a LOT of expert judgment about the causes and circumstances of the near misses. Hence, we were able with NATS experts to build a very detailed model that could be ‘validated’ against the actual near miss data. What is very interesting are what factors NATS needed in the model. The psychological state and stress of air traffic controllers was included in the model as were certain psychological traits of pilots. It turns out that certain airlines were more likely to be involved in a near-misses primarily because of traits of their pilots.

Tuesday, 24 March 2015

The problem with big data and machine learning


The advent of ‘big data’, coupled with fancy statistical machine learning techniques, is increasingly seducing people to believe that new insights and better predictions can be achieved in a wide range of important applications, without relying on the input of domain experts. The applications range from learning how to retain customers through to learning what makes people susceptible to particular diseases. I have written before about the dangers of this kind of 'learning' from data alone (no matter how 'big' the data is).

Contrary to the narrative being sold by the big data community, if you want accurate predictions and improved, decision-making then, invariably, you need to incorporate human knowledge and judgment. This enables you to build rational causal models based on 'smart' data. The main objections to using human knowledge - that it is subjective and difficult to acquire - are, of course,  key drivers of the big data movement. But this movement underestimates the typically very high costs of collecting, managing and analysing big data. So, the sub-optimal outputs you get from pure machine learning do not even come cheap.

To clarify the dangers of relying on big data and machine learning, and to show how smart data and causal modelling (using Bayesian networks) gives you better results, we have collected together the following short stories and examples:
The whole subject of 'smart data' rather than 'big data' is also the focus of the BAYES-KNOWLEDGE project.

Tuesday, 3 March 2015

The Statistics of Climate Change


From left to right: Norman Fenton, Hannah Fry, David Spiegelhalter. Link to the Programme's BBC website
Norman Fenton, 3 March 2015 (This is a cross posting of the article here)

I had the pleasure of being one of the three presenters of the BBC documentary called “Climate Change by Numbers”  (first) screened on BBC4 on 2 March 2015.

The motivation for the programme was to take a new look at the climate change debate by focusing on three key numbers that all come from the most recent IPCC report. The numbers were:
  • 0.85 degrees - the amount of warming the planet has undergone since 1880
  • 95% - the degree of certainty climate scientists have that at least half the warming in the last 60 years is man-made
  • one trillion tonnes - the cumulative amount of carbon that can be burnt, ever, if the planet is to stay below ‘dangerous levels’ of climate change
The idea was to get mathematicians/statisticians who had not been involved in the climate change debate to explain in lay terms how and why climate scientists had arrived at these three numbers. The other two presenters were Dr Hannah Fry (UCL) and Prof Sir David Spiegelhalter (Cambridge) and we were each assigned approximately 25 minutes on one of the numbers. My number was 95%.

Being neither a climate scientist nor a classical statistician (my research uses Bayesian probability rather than classical statistics to reason about uncertainty) I have to say that I found the complexity of the climate models and their underlying assumptions to be daunting. The relevant sections in the IPCC report are extremely difficult to understand and they use assumptions and techniques that are very different to the Bayesian approach I am used to. In our Bayesian approach we build causal models that combine prior expert knowledge with data. 

In attempting to understand and explain how the climate scientists had arrived at their 95% figure I used a football analogy – both because of my life-time interest in football and because - along with my colleagues Anthony Constantinou and Martin Neil – we have worked extensively on models for football prediction. The climate scientists had performed what is called an “attribution study” to understand the extent to which different factors – such as human CO2 emissions – contributed to changing temperatures. The football analogy was to understand the extent to which different factors contributed to changing success of premiership football teams as measured by the total number of points they achieved season-by-season.  In contrast to our normal Bayesian approach – but consistent with what the climate scientists did – we used data and classical statistical methods to generate a model of success in terms of the various factors. Unlike the climate models which involve thousands of variables we had to restrict ourselves to a very small number of variables (due to a combination of time limitations and lack of data). Specifically, for each team and each year we considered:
  • Wages (this was the single financial figure we used)
  • Total days of player injuries
  • Manager experience
  • Squad experience
  • Number of new players
The statistical model generated from these factors produced, for most teams, a good fit of success over the years for which we had the data. Our ‘attribution study’ showed wages was by far the major influence. When wages was removed from the study, the resulting statistical model was not a good fit. This was analogous to what the climate scientists’ models were showing when the human CO2 emissions factor was removed from their models; the previously good fit to temperature was no longer evident. And, analogous to the climate scientists’ 95% derived from their models, we were able to conclude there was a 95% chance that an increase in turnover of 10 per cent would result in at least one extra premiership point. (Update: note that this was a massive simplification to make the analogy. I am certainly not claiming that increasing wages causes an increase in points. If I had had the time I would have explained that in a proper model - like the Bayesian networks we have previously built - wages offered is one of the many factors influencing quality of players that can be bought which, in turn, along with other factors influences performance).

Obviously there was no time in the programme to explain either the details or the limitations of my hastily put-together football attribution study and I will no doubt receive criticism for it (I am preparing a detailed analysis).  But the programme also did not have the time or scope to address the complexity of some of the broader statistical issues involved in the climate debate (including issues that lead some climate scientists to claim the 95% figure is underestimated and others to believe it is overestimated). In particular, the issues that were not covered were:
  • The real probabilistic meaning of the 95% figure. In fact it comes from a classical hypothesis test in which observed data is used to test the credibility of the ‘null hypothesis’. The null hypothesis is the ‘opposite’ statement to the one believed to be true, i.e.  ‘Less than half the warming in the last 60 years is man-made’. If, as in this case, there is only a 5%  probability of observing the data if the null hypothesis is true, the statisticians equate this figure (called a p-value) to a 95% confidence that we can reject the null hypothesis. But the probability here is a statement about the data given the hypothesis. It is not generally the same as the probability of the hypothesis given the data (in fact equating the two is often referred to as the ‘prosecutors fallacy’, since it is an error often made by lawyers when interpreting statistical evidence).See here and here for more on the limitations of p-values and confidence intervals.
  • Any real details of the underlying statistical methods and assumptions. For example, there has been controversy about the way a method called principal component analysis was used to create the famous hockey stick graph that appeared in previous IPCC reports. Although the problems with that method were recognised it is not obvious how or if they have been avoided in the most recent analyses.
  •  Assumptions about the accuracy of historical temperatures. Much of the climate debate  (such as that concerning the exceptionalness of the recent rate of temperature increase) depends on assumptions about historical temperatures dating back thousands of years. There has been some debate about whether sufficiently large ranges were used.
  • Variety and choice of models. There are many common assumptions in all of the climate models used by the IPCC and it has been argued that there are alternative models not considered by the IPCC which provide an equally good fit to climate data, but which do not support the same conclusions.
Although I obviously have a bias, my enduring impression from working on the programme is that the scientific discussion about the statistics of climate change would benefit from a more extensive Bayesian approach. Recently some researchers have started to do this, but it is an area where I feel causal Bayesian network models could shed further light and this is something that I would strongly recommend.

Acknowledgements: I would like to thank the BBC team (especially Jonathan Renouf, Alex Freeman, Eileen Inkson, and Gwenan Edwards) for their professionalism, support, encouragement, and training; and my colleagues Martin Neil and Anthony Constantinou for their technical support and advice. 

My fee for presenting the programme has been donated to the charity Magen David Adom
Watching the programme as it is screened

Saturday, 15 November 2014

How to measure anything

Douglas Hubbard (left) and Norman Fenton in London 15 Nov 2014.
If you want to know how to use measurement to reduce risk and uncertainty in a wide range of business applications, then there is no better book than Douglas Hubbard's "How to Measure Anything: Finding the Value of Intangibles in Business" (now in its 3rd edition). Douglas is also the author of the excellent "The Failure of Risk Management: Why It's Broken and How to Fix It".

Anyone who has read our Bayesian Networks book or the latest (3rd edition) of my Software Metrics book (the one I gave Douglas in the above picture!) will know how much his work has influenced us recently.

Although we have previously communicated about technical issues by email, today I had the pleasure of meeting Douglas for the first time when we were able to meet for lunch in London.We discussed numerous topics of mutual interest (including the problems with classical hypothesis testing - and how Bayes provides a better alternative, and evolving work on the 'value of information' which enables you to identify where to focus your measurement to optimise your decision-making).

Tuesday, 17 June 2014

Proving referee bias with Bayesian networks

An article in today's Huffington Post by Raj Persaud and Adrian Furnham talks about the scientific evidence that supports the idea of referee bias in football. One of the studies they describe is the recent work by Anthony Constantinou, Norman Fenton and Liam Pollock** which developed a causal Bayesian network model to determine referee bias and applied it to the data from all matches played in the 2011-12 Premier League season. Here is what they say about our study:
Another recent study might just have scientifically confirmed this possible 'Ferguson Factor', entitled, 'Bayesian networks for unbiased assessment of referee bias in Association Football'. The term 'Bayesian networks', refers to a particular statistical technique deployed in this research, which mathematically analysed referee bias with respect to fouls and penalty kicks awarded during the 2011-12 English Premier League season.
The authors of the study, Anthony Constantinou, Norman Fenton and Liam Pollock found fairly strong referee bias, based on penalty kicks awarded, in favour of certain teams when playing at home.
Specifically, the two teams (Manchester City and Manchester United) who finished first and second in the league, appear to have benefited from bias that cannot be explained by other factors. For example a team may be awarded more penalties simply because it's more attacking, not just because referees are biased in its favour.

The authors from Queen Mary University of London, argue that if the home team is more in control of the ball, then, compared to opponents, it's bound to be awarded more penalties, with less yellow and red cards, compared to opponents. Greater possession leads any team being on the receiving end of more tackles. A higher proportion of these tackles are bound to be committed nearer to the opponent's goal, as greater possession also usually results in territorial advantage.
However, this study, published in the academic journal 'Psychology of Sport and Exercise', found, even allowing for these other possible factors, Manchester United with 9 penalties awarded during that season, was ranked 1st in positive referee bias, while Manchester City with 8 penalties awarded is ranked 2nd. In other words it looks like certain teams (most specifically Manchester United) benefited from referee bias in their favour during Home games, which cannot be explained by any other possible element of 'Home Advantage'. 
What makes this result particularly interesting, the authors argue, is that for most of the season, these were the only two teams fighting for the English Premiere League title. Were referees influenced by this, and it impacted on their decision-making?  Conversely the study found Arsenal, a team of similar popularity and wealth, and who finished third, benefited least of all 20 teams from referee bias at home, with respect to penalty kicks awarded. With the second largest average attendance as well as the second largest average crowd density, Arsenal were still ranked last in terms of referee bias favouring them for penalties awarded. In other words, Arsenal didn't seem to benefit much at all from the kind of referee bias that other teams were gaining from 'Home Advantage'. Psychologists might argue that temperament-wise, Sir Alex Ferguson and Arsene Wenger appear at opposite poles of the spectrum.
**  Constantinou, A. C., Fenton, N. E., & Pollock, L. (2014). "Bayesian networks for unbiased assessment of referee bias in Association Football". To appear in Psychology of Sport & Exercise. A pre-publication draft can be found here.

Our related work on using Bayesian networks to predict football results is discussed here.

Wednesday, 2 April 2014

Bayesian network approach to Drug Economics Decision Making


Consider the following problem:
A relatively cheap drug (drug A) has been used for many years to treat patients with disease X. The drug is considered quite successful since data reveals that 85% of patients using it have a ‘good outcome’ which means they survive for at least 2 years. The drug is also quite cheap, costing on average $100 for a prolonged course. The overall “financial benefit” of the drug (which assumes a ‘good outcome’ is worth $5000 and is defined as this figure minus the cost) has a mean of $4985.

There is an alternative drug (drug B) that a number of specialists in disease X strongly recommend. However, the data reveals that only 65% of patients using drug B survive for at least 2 years (Fig. 1(b)). Moreover, the average cost of a prolonged course is $500. The overall “financial benefit” of the drug has a mean of just $2777.
On seeing the data the Health Authority recommends a ban against the use of drug B. Is this a rational decision?

The answer turns out to be no. The short paper here explains this using a simple Bayesian network model that you can run (by downloading the free copy of AgenaRisk)