A nice 2-page article about our BAYES-KNOWLEDGE project is in the latest issue of EU Research Magazine Beyond the Horizon. A pdf version is here.
Showing posts with label Bayes Knowledge. Show all posts
Showing posts with label Bayes Knowledge. Show all posts
Wednesday, 14 March 2018
Wednesday, 1 June 2016
Bayesian networks for Cost, Benefit and Risk Analysis of Agricultural Development Projects
Successful implementation of major projects requires careful management of uncertainty and risk. Yet, uncertainty is rarely effectively calculated when analysing project costs and benefits. In the case of major agricultural and other development projects in Africa this challenge is especially important.
A paper just published* in the journal Experts Systems with Applications presents a Bayesian network (BN) modelling framework to calculate the costs, benefits, and return on investment of a project over a specified time period, allowing for changing circumstances and trade-offs. Marianne Gadeberg and Eike Luedeling have written an overview of the work here.
The framework uses hybrid and dynamic BNs containing both discrete and continuous variables over multiple time stages. The BN framework calculates costs and benefits based on multiple causal factors including the effects of individual risk factors, budget deficits, and time value discounting, taking account of the parameter uncertainty of all continuous variables. The framework can serve as the basis for various project management assessments and is illustrated using a case study of an agricultural development project. The work was a collaboration between the World Agroforestry Centre (ICRAF), Nairobi, Kenya, the Risk Information Management Group at Queen Mary (as part of the BAYES-KNOWLEDGE project) and Agena Ltd.
*The full reference is:
Yet, B., Constantinou, A., Fenton, N., Neil, M., Luedeling, E., & Shepherd, K. (2016). "A Bayesian Network Framework for Project Cost, Benefit and Risk Analysis with an Agricultural Development Case Study" . Expert Systems with Applications, Volume 60, 30 October 2016, Pages 141–155. DOI: 10.1016/j.eswa.2016.05.005.Until July 2016 the full published pdf is available for free. A permanent pre-publication pdf is available here.
See also: Can we build a better project: assessing complexities in development projects
Acknowledgements: Part of this work was performed under the auspices of EU project ERC-2013-AdG339182-BAYES_KNOWLEDGE and part under ICRAF Contract No SD4/2012/214 issued to Agena. We acknowledge support from the Water, Land and Ecosystems (WLE) program of the Consultative Group on International Agricultural Research (CGIAR).
Thursday, 26 May 2016
Using Bayesian networks to assess new forensic evidence in an appeal case
If new forensic evidence becomes available after a conviction how do lawyers determine whether it raises sufficient questions about the verdict in order to launch an appeal? It turns out that there is no systematic framework to help lawyers do this. But a paper published today by Nadine Smit and colleagues in Crime Science presents such a framework driven by a recent case, in which a defendant was convicted primarily on the basis of sound evidence, but where subsequent analysis of the evidence revealed additional sounds that were not considered during the trial.
From the case documentation, we know the following:
- A baby was injured during an incident on the top floor of a house
- Blood from the baby was found on the wall in one of the rooms upstairs
- On an audio recording of the emergency telephone call made by the suspect, a scraping sound (allegedly indicating scraping blood off a wall) can be heard
- The suspect was charged with attempted murder
The framework described in Smit's paper is intended to overcome the gap between what is generally known from scientific analyses and what is hypothesized in a legal setting. It is based on Bayesian networks (BNs) which are a structured and understandable way to evaluate the evidence in the specific case context and present it in a clear manner in court. However, BN methods are often criticised for not being sufficiently transparent for legal professionals. To address this concern the paper shows the extent to which the reasoning and decisions of the particular case can be made explicit and transparent. The BN approach enables us to clearly define the relevant propositions and evidence, and uses sensitivity analysis to assess the impact of the evidence under different prior assumptions. The results show that such a framework is suitable to identify information that is currently missing, and clearly crucial for a valid and complete reasoning process. Furthermore, a method is provided whereby BNs can serve as a guide to not only reason with incomplete evidence in forensic cases, but also identify very specific research questions that should be addressed to extend the evidence base to solve similar issues in the future.
Full citation:
Smit, N. M., Lagnado, D. A., Morgan, R. M., & Fenton, N. E. (2016). "An investigation of the application of Bayesian networks to case assessment in an appeal case". Crime Science, 2016, 5: 9, DOI 10.1186/s40163-016-0057-6 (open source). Published version pdf.The research was funded by the Engineering and Physical Sciences Research Council of the UK through the Security Science Doctoral Research Training Centre (UCL SECReT) based at University College London (EP/G037264/1), and the European Research Council (ERC-2013-AdG339182-BAYES_KNOWLEDGE).
The BN model (which is fully spceified in the paper) was built and run using the free version of AgenaRisk.
Tuesday, 8 March 2016
Anthony Constantinou presents BAYES-KNOWLEDGE work in Parliament
![]() |
| Anthony Constantinou and his poster |
Updated 8 March 2016
Anthony Constantinou was selected to present his work (summarised in this poster) on the BAYES-KNOWLEDGE project in Parliament on 7 March 2016 as part of the SET for Britain awards 2016 This was the Press Release from Parliament:
Dr Anthony Constantinou, 31, a Post-Doctoral Researcher at Queen Mary University of London, hailing from Limassol, Cyprus, is attending Parliament to present his mathematics research to a range of politicians and a panel of expert judges, as part of SET for Britain on Monday 7 March.Although Anthony did not win one of the final prizes his entry was highly praised by the judges and many MPs who he spoke with.
Anthony’s poster on research about Decision Systems which are based on probabilistic graphical models for uncertainty quantification and risk management, will be judged against dozens of other scientists’ research in the only national competition of its kind. Anthony is funded as part of the BAYES-KNOWLEDGE project (http://bayes-knowledge.com/).
Anthony was shortlisted from hundreds of applicants to appear in Parliament.
On presenting his research in Parliament, he said, “This is a great opportunity to demonstrate the art of intelligent decision making, which is merely based on identifying the best possible set of actions to be taken on the basis of trade-offs in an effort to achieve some objectives; then having to explain why an objective has not been met! Having applied these methods to a diverse range of real-world application domains, spanning from forensics and medical sciences to sports and gambling market efficiency, I do hope I can keep the audience entertained.”
Stephen Metcalfe MP, Chairman of the Parliamentary and Scientific Committee, said:
“This annual competition is an important date in the parliamentary calendar because it gives MPs an opportunity to speak to a wide range of the country’s best young researchers.
“These early career engineers, mathematicians and scientists are the architects of our future and SET for Britain is politicians’ best opportunity to meet them and understand their work.”
Anthony’s research has been entered into the Mathematical Sciences session of the competition, which will end in a gold, silver and bronze prize-giving ceremony.
Judged by leading academics, the gold medalist receives £3,000, while silver and bronze receive £2,000 and £1,000 respectively.
The Parliamentary and Scientific Committee runs the event in collaboration with the Royal Academy of Engineering, the Royal Society of Chemistry, the Institute of Physics, the Royal Society of Biology, The Physiological Society and the Council for the Mathematical Sciences, with financial support from Essar, the Clay Mathematics Institute, Warwick Manufacturing Group (WMG), the Institute of Biomedical Science, the Bank of England and the Society of Chemical Industry.
| Anthony about to enter the Houses of Parliament |
| Anthony with Nick Clegg and David Cameron! |
| Anthony with poster |
![]() |
| The prize giving bit with Science Minister Joe Johnson (5th from right) |
Improving Bayesian networks by learning from similar ones
A major challenge of the BAYES-KNOWLEDGE project is about how to build useful and accurate Bayesian network (BN) models for decision support when there is little relevant data. Much of what we have done involves exploiting expert judgment. But my colleagues Yun Zhou and Tim Hospedales have developed a so-called 'tranfer learning' method that enables us to leverage data from different but related problems. Suppose, for example, we have a BN model for a particular medical diagnostic problem that we built based on limited data and expert judgment in the UK. But suppose also that a model for the same (or very similar) diagnostic problem has been developed in the USA based on a much larger data set. Some assumptions in the US model will be different to the UK (such as the population demographics or particular testing methods) but much of the underlying pathology will be the same. The challenge is to understand and exploit the heterogeneous relatedness of the models.
The result of this work has just been published in an article in the Elsevier journal Expert Systems with Applications. Elsevier have provided free access to the article until April 24, 2016:
http://authors.elsevier.com/a/1SfM-3PiGT01St
The article describes a new transfer learning algorithm for improved BN parameter learning, and the experimental results demonstrate its superiority compared to other state-of-the-art parameter transfer methods. The method is applied to a real-world medical case study, namely the problem of trauma care (a problem for which our team had initially developed a decision support BN model in collaboration with UK medics).
Full reference for the new article:
Zhou, Y., Hospedales, T., Fenton, N. E. (2016), "When and where to transfer for Bayes net parameter learning", Expert Systems with Applications. 55, 361-373
http://dx.doi.org/10.1016/j.eswa.2016.02.011.
Sunday, 18 October 2015
What is the value of missing information when assessing decisions that involve actions for intervention?
This is a summary of the following new paper:
Most decision support models in the medical domain provide a prediction about a single key unknown variable, such as whether a patient exhibiting certain symptoms is likely to have (or develop) a particular disease.
However we seek to enhance decision analysis by determining whether a decision based on such a prediction could be subject to amendments on the basis of some incomplete information within the model, and whether it would be worthwhile for the decision maker to seek further information prior to the decision. In particular we wish to incorporate interventional actions and counterfactual analysis, where:

The type of predictive decision support models to which our work applies are Bayesian networks. These are graphical models which represent the causal or influential relationships between a set of variables and which provide probabilities for each unknown variable.
The method is applied to two real-world Bayesian network models that were previously developed for decision support in forensic medical sciences. In these models a decision maker (such as a probation officer or a clinician) has to determine whether to release a prisoner/patient based on the probability of the (unknown) hypothesis variable: “individual violently reoffends after release”. Prior to deciding on release, the decision maker has the option to simulate various interventions to determine whether an individual’s risk of violence can be managed to acceptable levels. Additionally, the decision maker may have the option to gather further information about the individual. It is possible that knowing one or more of these unobserved factors may lead to a different decision about release.
We used the method to examine the average information gain; that is, what we learn about the importance of the factors that remain unknown within the model. Based on six different sets of experiments with various assumptions we show that:
The method becomes useful for decision makers, not only when decision making is subject to amendments on the basis of some unknown risk factors, but also when it is not. Knowing that a decision outcome is independent of one or more unknown risk factors saves us from seeking information about that particular set of risk factors.
This summary can also be found on the Atlas of Science
Constantinou AC, Yet B, Fenton N, Neil M, Marsh W "Value of Information analysis for interventional and counterfactual Bayesian networks in forensic medical sciences". Artif Intell Med. 2015 Sep 8 doi:10.1016/j.artmed.2015.09.002. The full pre-publication version can be found here.
Most decision support models in the medical domain provide a prediction about a single key unknown variable, such as whether a patient exhibiting certain symptoms is likely to have (or develop) a particular disease.
However we seek to enhance decision analysis by determining whether a decision based on such a prediction could be subject to amendments on the basis of some incomplete information within the model, and whether it would be worthwhile for the decision maker to seek further information prior to the decision. In particular we wish to incorporate interventional actions and counterfactual analysis, where:
- An interventional action is one that can be performed to manipulate the effect of some desirable future outcome. In medical decision analysis, an intervention is typically represented by some treatment, which can affect a patient’s health outcome.
- Counterfactual analysis enables decision makers to compare the observed results in the real world to those of a hypothetical world; what actually happened and what would have happened under some different scenario.
The type of predictive decision support models to which our work applies are Bayesian networks. These are graphical models which represent the causal or influential relationships between a set of variables and which provide probabilities for each unknown variable.
The method is applied to two real-world Bayesian network models that were previously developed for decision support in forensic medical sciences. In these models a decision maker (such as a probation officer or a clinician) has to determine whether to release a prisoner/patient based on the probability of the (unknown) hypothesis variable: “individual violently reoffends after release”. Prior to deciding on release, the decision maker has the option to simulate various interventions to determine whether an individual’s risk of violence can be managed to acceptable levels. Additionally, the decision maker may have the option to gather further information about the individual. It is possible that knowing one or more of these unobserved factors may lead to a different decision about release.
We used the method to examine the average information gain; that is, what we learn about the importance of the factors that remain unknown within the model. Based on six different sets of experiments with various assumptions we show that:
- the average relative percentage gain in terms of Value of Information ranged between 11.45% and 59.91% (where a gain of X% indicates an expected X% relative reduction of the risk of violent reoffence);
- the potential amendments in Decision Making, as a result of the expected information gain, ranged from 0% to 86.8% (where an amendment of X% indicates that X% of the initial decisions are expected to have been altered).
The method becomes useful for decision makers, not only when decision making is subject to amendments on the basis of some unknown risk factors, but also when it is not. Knowing that a decision outcome is independent of one or more unknown risk factors saves us from seeking information about that particular set of risk factors.
This summary can also be found on the Atlas of Science
Thursday, 15 October 2015
Talk: Bayesian networks: why smart data is better than big data
by Prof. Norman Fenton from the School of Electronic
Engineering and Computer Science (QMUL)
WHEN: Fri, 16th October 2 - 3 pm
WHERE: People's Palace PP2 (Mile End Campus)
"This talk will
provide an introduction to Bayesian networks which, due to relatively recent
algorithmic breakthroughs, has become an increasingly popular technique for
risk assessment and decision analysis. I will provide an overview of successful
applications (including transport safety, medical, law/forensics, operational
risk, and football prediction). What is common to all of these applications is
that the Bayesian network models are built using a combination of expert
judgment and (often very limited) data. I will explain why Bayesian networks
‘learnt’ purely from data – even when ‘big data’ is available - generally do
not work well."
All are welcome. The
seminar consists of an app. 45 min long lecture and discussion.
In case of any
questions, feel free to contact me.
Hope to see you
tomorrow,
Judit Petervari
____________________
Judit Petervari
PhD Student
Biological and Experimental Psychology Group
School of Biological and Chemical Sciences
Queen Mary University of London
Mile End Road
E1 4NS London
United Kingdom
Biological and Experimental Psychology Group
School of Biological and Chemical Sciences
Queen Mary University of London
Mile End Road
E1 4NS London
United Kingdom
E-mail: j.petervari@qmul.ac.uk
Office: G.E. Fogg Building, Room 2.16
Office: G.E. Fogg Building, Room 2.16
Tuesday, 15 September 2015
Yet another flawed statistical study attracts massive unquestioning attention
![]() |
| The Guardian, 29 Sept 2015 |
“The research clearly shows what we have been talking about for a while: that diversity leads to better decision-making”.As is typical when the results of a statistical study fit a popular narrative, the story attracted massive, unquestioning attention. Unfortunately, while I am sure that most people agree that greater gender diversity in the Boardroom is a worthy objective, based on the ‘full report’ – and in the absence of other data - Lagerberg's claim is simply not supported. In fact, the study exemplifies some of the classic misuses of statistics that we wrote about in the first chapter of our book and highlights yet again the need for proper causal/explanatory models to be used in statistical studies such as these*.
Moreover, using the data in Lagerberg's study it is possible to construct a simple causal model (a Bayesian network) that replicates the results but with provably opposite conclusions: diversity decreases performance.
The full report and BN model are provided here. The model can be run in the free version of AgenaRisk.
*Making such an approach both universally feasible and acceptable is the major objective of BAYES-KNOWLEDGE.
Sunday, 30 August 2015
Using Bayesian networks to assess and manage risk of violent reoffending among prisoners
![]() |
| Fragment of BN model |
The (BN) model, called DSVM-P (Decision Support for Violence Management – Prisoners) captures the causal relationships between risk factors, interventions and violence. It also allows for specific risk factors to be targeted for causal intervention for risk management of future re-offending. These decision support features are not available in the previous generation of models used by practitioners and forensic psychiatrists.
Full reference:
Constantinou, A., Freestone M., Marsh, W., Fenton, N. E. , Coid, J. (2015) "Risk assessment and risk management of violent reoffending among prisoners", Expert Systems With Applications 42 (21), 7511-7529. Published version: http://dx.doi.org/10.1016/j.eswa.2015.05.025.
Download Pre-publication draft.
Friday, 17 July 2015
The use of Bayes in the Netherlands Appeal Court
![]() |
| Henry Prakken |
There has been an important development on the use of Bayes in the Law in the Netherlands, with what is possibly the first full Bayesian analysis of a major crime in an appeal court there.
The case, referred to as the “Breda 6”, was the 1993 murder of a Chinese woman in Breda in her son’s restaurant. Six young people were convicted of the crime and sentenced to up to 10 years in jail (all have since completed their sentences). In 2012 the advocate general recommended the case be looked at again because it centred on confessions which may have been false.
In the review of the case a Bayesian argument supporting the prosecution case was presented by Frans Alkemade (Update: see below about some concerns Frans has about this article). Frans is actually a physicist who previously used Bayes to analyse a drugs trafficking case, concluding in a report commissioned by the prosecution, that there was not enough evidence for a conviction (the suspect was acquitted). The court requested that Henry Prakken (professor in Legal Informatics and Legal Argumentation at the University of Groningen) respond to the Bayesian argument.
In June, while he was preparing his response, I met Henry at the International Conference on AI and the Law in San Diego. Henry told me about the case and Alkemade's analysis for which the guilty hypothesis was "At least some of the six suspects were involved in the crime, which took place after 4:30 on the night of 3 and 4 july 1993, and which included the luring of the victim to the restaurant by at least some of the female suspects". Alkemade interpreted "involved" in a weak way and Henry said:
"..it could even be no more than just knowing about the crime. In fact, one of my points of criticism was that this guilt hypothesis is not useful for the court, since it is consistent with the innocence of any of the individual suspects (and even with the collective innocence of all three male suspects)."Among other things, Alkemade focused on two pieces of evidence:
- A report by the Criminal Intelligence Unit (CID) of the Dutch police, saying that they had received information from "usually reliable" sources identifying two of the male defendants and one of the female defendants as being involved in the murder**.
- The subsequent discovery that two of the female defendants (not mentioned in the CID report and who supposedly knew the three defendants mentioned in the CID report) worked next door to the murder scene.
The intention was not to replicate all of the features of the case but rather to demonstrate the impact of missing dependencies in Alkemade's argument. Indeed, with a range of reasonable assumptions, the BN model pointed to a much lower probability of guilt than suggested by Alkemade's calculations.
Henry presented his response in court last week. He said:
"The court session was sometimes frustrating, since the discussion was fragmentary and sometimes I had the impression that the court felt it was confronted with a battle of the experts without the means to understand who was right."In my view this case confirms our claim that presenting a Bayesian legal argument from first principles (as Alkemade did) is not a good idea. The very fact that people assume it is necessary to do this for Bayes to be accepted is actually the reason (ironically) that there will continue to be very strong resistance to accepting Bayes in the courtroom. Why? Because it means you are restricted to ludicrously over-simplistic (and normally flawed) models of the case (3 unknowns maximum) because that is the limit of the Bayesian calculations you can do by hand and explain clearly. Our proposed solution is to model (and run) the case properly using a BN and report back on the results stating in lay terms what the model assumptions were and how sensitive the conclusions are to different prior assumptions.
**Henry told me that there was something funny with the CID report, as also noted by the advocate general. According to the CID report, the anonymous informant had also accused the defendants mentioned in the report of committing several other crimes, but in the investigations preceding the revision case the police investigators had not been able to find any confirmation of these other crimes, not even reports of these supposed crimes to the police by the supposed victims. As the advocate general stated in 2012, this casts doubt on the reliability of the CID informants.
***My colleague Richard Gill - who knows Alkemade - says that Alkemade was careful to define his pieces of evidence in such a way that he thinks that he can justify the independence assumptions which he needs in order to at least conservatively bound the Likelihood ratio coming from each piece of evidence in turn.
UPDATE 21 July 2015: Frans has contacted me stating a number of concerns about the above narrative and provided a number of technical insights that I was not aware of. As an expert witness in a case that is still under trial, he does not feel free to discuss any details in public, but once the trial has finished I will provide an updated report that incorporates his comments. What I can confirm, however, is that in order to do the calculations manually Frans could not model dependencies between different pieces of evidence - a major limitation - although he did make clear the limitations and pitfalls of this in his report.
See also:
- Review of the use of Bayes in the Law (pdf report)
- A General Structure for Legal Arguments About Evidence Using Bayesian Networks (published article)
- Assessing evidence and testing appropriate hypotheses (published article)
- Barry George case: new insights on the evidence.
- Sally Clark case: another statistical oversight
- Ben Geen: another possible case of miscarriage of justice and misunderstanding of statistics
Thursday, 9 July 2015
Why target setting leads to poor decision-making
Norman Fenton is the co-author of an article in Nature published today that addresses the issue of improved decision-making in the context of international sustainable development goals. The article pushes for a Bayesian, smart-data approach:
We contend that target-setting is flawed, costly and could have little — or even negative — impact. First, targets may have unintended consequences. For example, education quality as a whole suffered in some countries that diverted resources to early schooling to meet the target of the Millennium Development Goal (MDG) of achieving universal primary education.The approach is based on five principles:
Second, target-setting inhibits learning by focusing efforts on meeting the target rather than solving the problem. The milestones are easily manipulated — aims such as halving deaths from road-traffic accidents can trigger misreporting if the performance falls short or encourage underperformance if the goal can be exceeded.
Third, it is costly: development partners will have to reallocate scant resources for a 'data revolution' that will cost an estimated US$1 billion a year.
We advocate a different approach. Governments and the development community need to embrace decision-analysis concepts and tools that have been used for decades in mining, oil, cybersecurity, insurance, environmental policy and drug development.
- Replace targets with measures of investment return
- Model intervention decisions
- Integrate expert knowledge
- Include uncertainty in predictive models
- Measure the most informative variables
It is a common mistake to assume that 'evidence' is the same as 'data' or that 'subjective' means 'uninformative'. Decision-making should draw on all appropriate sources of evidence. In developing countries where data are sparse, expert knowledge can fill the gaps. For instance, in our assessment of the viability of agroforestry projects in Africa, we used our experience to set ranges on tree-survival rates, costs of raising tree seedlings and farm prices of tree products.
....
Decision theorists and local experts will have to work together to identify relevant variables, causal associations and uncertainties. The most widely accepted method of incorporating knowledge for probability assessment is Bayes' theorem. This updates the likelihood of a belief in some event (such as whether an intervention will reduce poverty) when observing new evidence about the event (such as the occurrence of drought). Bayesian analyses — incorporating historical data and expert judgement — are used in transport and systems-safety assessments, medical diagnosis, operational risk assessment in finance and in forensics, but seldom in development. They should be used, for example, to evaluate the relative risks of competing development interventions.
....
Decision-makers .. should employ probabilistic decision analysis, for example Monte Carlo simulations or Bayesian network models. Provided that such models are developed using properly calibrated expert judgement and decision-focused data, they can incorporate the key factors and outcomes and the causal relationships between them. For instance, simulations for evaluating options for building a water pipeline could take into account rare 'what-if' scenarios, such as a hurricane during development, and predict (with probabilities) the time and cost of implementation and the benefits of improved water supply.
Thursday, 26 March 2015
The risk of flying
Norman Fenton, 26 March 2015
I have just done an interview on BBC Radio Scotland about aircraft safety in the light of the GermanWings crash - which now appears to have been a deliberate act of sabotage by the co-pilot*. I have uploaded a (not very good) recording of it here (mp3 file - it is just under 4 minutes) or here (a more compact m4a file)
Because this type of event is so rare classical frequentist statistics provides no real help when it comes to risk assessment. In fact, it is exactly the kind of risk assessment problem for which you need causal models and expert judgement (as explained in our book) if you want any kind of risk insights.
Irrespective of this particular incident, the interview gave me the opportunity to highlight a very common myth, namely that “flying is the safest form of travel”. If you look at deaths per million travellers then, indeed, there are 50 times as many car deaths as plane deaths. However, this is a silly measure because there are so many more car travellers than plane travellers. So, typically, analysts use deaths per million miles travelled; with respect to this measure car travel is still 'riskier' than air travel, but the death rate is only about twice as high as plane deaths. But this measure is also biased in favour of planes because the average plane journey is much further than the average car journey.
So a much fairer measure is the number of deaths per passenger journey. And for this, the rate of plane deaths is actually three times higher than car deaths; in fact only bikes and motorbikes are worse than planes.
Despite all this there is still a very low probability of a plane journey resulting in fatalities - about 1 in half a million (and much less on commercial flights in Western Europe). However, if we have reason to believe that, say, recent converts to a terrorist ideology have been training and becoming pilots then the probability of the next plane journey resulting in fatalities becomes much higher, despite the past data.
*I had an hour’s notice of the interview and was told what I would be asked. I was actually not expecting to be asked about how to assess the risk of this specific type of incident; I was assuming I would only be asked about aircraft safety risk in general and about the safety record of the A320.
Postscript: Following the interview a colleage asked:
"Did you have the mental issues of the co-pilot on the radar when you replied? "
My response: Interesting question. A few years back we were involved extensively in work with NATS (National Air Traffic Safety) to model/predict risk of mid-air collision over the UK airspace. In particular NATS wanted to know how the probability of a mid-air collision might change given different proposals for changes to the ATM architecture (e.g. ‘adding new ground radar stations’ versus ‘adding new on-board collisions alert systems’). Now - apart from three incidents in the late 1940’s which all involved at least one military jet - there has not been any actual mid-air collisions over UK airspace (so negligible data there) and the proposed technology was ‘new’ (so no directly relevant data there) but there was a LOT of data on "near misses" of different degrees of seriousness and a LOT of expert judgment about the causes and circumstances of the near misses. Hence, we were able with NATS experts to build a very detailed model that could be ‘validated’ against the actual near miss data. What is very interesting are what factors NATS needed in the model. The psychological state and stress of air traffic controllers was included in the model as were certain psychological traits of pilots. It turns out that certain airlines were more likely to be involved in a near-misses primarily because of traits of their pilots.
Tuesday, 24 March 2015
The problem with big data and machine learning
The advent of ‘big data’, coupled with fancy statistical machine learning techniques, is increasingly seducing people to believe that new insights and better predictions can be achieved in a wide range of important applications, without relying on the input of domain experts. The applications range from learning how to retain customers through to learning what makes people susceptible to particular diseases. I have written before about the dangers of this kind of 'learning' from data alone (no matter how 'big' the data is).
Contrary to the narrative being sold by the big data community, if you want accurate predictions and improved, decision-making then, invariably, you need to incorporate human knowledge and judgment. This enables you to build rational causal models based on 'smart' data. The main objections to using human knowledge - that it is subjective and difficult to acquire - are, of course, key drivers of the big data movement. But this movement underestimates the typically very high costs of collecting, managing and analysing big data. So, the sub-optimal outputs you get from pure machine learning do not even come cheap.
To clarify the dangers of relying on big data and machine learning, and to show how smart data and causal modelling (using Bayesian networks) gives you better results, we have collected together the following short stories and examples:
- A short story illustrating why pure machine learning (without expert input) may be doomed to fail and totally unnecessary (2 page pdf)
- Another machine learning fable: explains why pure machine learning for identifying credit risk may result in perfectly incorrect risk assessment (1 page pdf)
- Moving from big data and machine learning to smart data and causal modelling: a simple example from consumer research and marketing (7 page pdf)
- A Bayesian Network for a simple example of Drug Economics Decision Making (4 page pdf)
Tuesday, 3 March 2015
The Statistics of Climate Change
![]() |
| From left to right: Norman Fenton, Hannah Fry, David Spiegelhalter. Link to the Programme's BBC website |
I had the pleasure of being one of the three presenters of the BBC documentary called “Climate Change by Numbers” (first) screened on BBC4 on 2 March 2015.
The motivation for the programme was to take a new look at
the climate change debate by focusing on three key numbers that all come from
the most recent IPCC report. The numbers were:
- 0.85 degrees - the amount of warming the planet has undergone since 1880
- 95% - the degree of certainty climate scientists have that at least half the warming in the last 60 years is man-made
- one trillion tonnes - the cumulative amount of carbon that can be burnt, ever, if the planet is to stay below ‘dangerous levels’ of climate change
The idea was to get mathematicians/statisticians who had not
been involved in the climate change debate to explain in lay terms how and why
climate scientists had arrived at these three numbers. The other two presenters
were Dr Hannah Fry (UCL) and Prof Sir David Spiegelhalter (Cambridge) and we
were each assigned approximately 25 minutes on one of the numbers. My number
was 95%.
Being neither a climate scientist nor a classical
statistician (my research uses Bayesian probability rather than classical statistics to reason about uncertainty) I have to say that I found the
complexity of the climate models and their underlying assumptions to be
daunting. The relevant sections in the IPCC report are extremely difficult to
understand and they use assumptions and techniques that are very different to
the Bayesian approach I am used to. In our Bayesian approach we build causal
models that combine prior expert knowledge with data.
In attempting to
understand and explain how the climate scientists had arrived at their 95%
figure I used a football analogy – both because of my life-time interest in
football and because - along with my colleagues Anthony Constantinou and Martin
Neil – we have worked extensively on models for football prediction. The
climate scientists had performed what is called an “attribution study” to
understand the extent to which different factors – such as human CO2 emissions
– contributed to changing temperatures. The football analogy was to understand
the extent to which different factors contributed to changing success of
premiership football teams as measured by the total number of points they
achieved season-by-season. In contrast
to our normal Bayesian approach – but consistent with what the climate
scientists did – we used data and classical statistical methods to generate a
model of success in terms of the various factors. Unlike the climate models
which involve thousands of variables we had to restrict ourselves to a very
small number of variables (due to a combination of time limitations and lack of
data). Specifically, for each team and each year we considered:
- Wages (this was the single financial figure we used)
- Total days of player injuries
- Manager experience
- Squad experience
- Number of new players
Obviously there was no time in the programme to explain either the details or the limitations of my hastily put-together football attribution study and I will no doubt receive criticism for it (I am preparing a detailed analysis). But the programme also did not have the time or scope to address the complexity of some of the broader statistical issues involved in the climate debate (including issues that lead some climate scientists to claim the 95% figure is underestimated and others to believe it is overestimated). In particular, the issues that were not covered were:
- The real probabilistic meaning of the 95% figure. In fact it comes from a classical hypothesis test in which observed data is used to test the credibility of the ‘null hypothesis’. The null hypothesis is the ‘opposite’ statement to the one believed to be true, i.e. ‘Less than half the warming in the last 60 years is man-made’. If, as in this case, there is only a 5% probability of observing the data if the null hypothesis is true, the statisticians equate this figure (called a p-value) to a 95% confidence that we can reject the null hypothesis. But the probability here is a statement about the data given the hypothesis. It is not generally the same as the probability of the hypothesis given the data (in fact equating the two is often referred to as the ‘prosecutors fallacy’, since it is an error often made by lawyers when interpreting statistical evidence).See here and here for more on the limitations of p-values and confidence intervals.
- Any real details of the underlying statistical methods and assumptions. For example, there has been controversy about the way a method called principal component analysis was used to create the famous hockey stick graph that appeared in previous IPCC reports. Although the problems with that method were recognised it is not obvious how or if they have been avoided in the most recent analyses.
- Assumptions about the accuracy of historical temperatures. Much of the climate debate (such as that concerning the exceptionalness of the recent rate of temperature increase) depends on assumptions about historical temperatures dating back thousands of years. There has been some debate about whether sufficiently large ranges were used.
- Variety and choice of models. There are many common assumptions in all of the climate models used by the IPCC and it has been argued that there are alternative models not considered by the IPCC which provide an equally good fit to climate data, but which do not support the same conclusions.
Although I obviously have a bias, my enduring impression
from working on the programme is that the scientific discussion about the
statistics of climate change would benefit from a more extensive Bayesian
approach. Recently some researchers have started to do this, but it is an area
where I feel causal Bayesian network models could shed further light and this
is something that I would strongly recommend.
Acknowledgements:
I would like to thank the BBC team (especially Jonathan Renouf, Alex Freeman,
Eileen Inkson, and Gwenan Edwards) for their professionalism, support,
encouragement, and training; and my colleagues Martin Neil and Anthony
Constantinou for their technical support and advice.
My fee for presenting the programme has been donated to the charity Magen David Adom
![]() |
| Watching the programme as it is screened |
Saturday, 15 November 2014
How to measure anything
| Douglas Hubbard (left) and Norman Fenton in London 15 Nov 2014. |
Anyone who has read our Bayesian Networks book or the latest (3rd edition) of my Software Metrics book (the one I gave Douglas in the above picture!) will know how much his work has influenced us recently.
Although we have previously communicated about technical issues by email, today I had the pleasure of meeting Douglas for the first time when we were able to meet for lunch in London.We discussed numerous topics of mutual interest (including the problems with classical hypothesis testing - and how Bayes provides a better alternative, and evolving work on the 'value of information' which enables you to identify where to focus your measurement to optimise your decision-making).
Tuesday, 17 June 2014
Proving referee bias with Bayesian networks
An article in today's Huffington Post by Raj Persaud and Adrian Furnham
talks about the scientific evidence that supports the idea of referee
bias in football. One of the studies they describe is the recent work by Anthony Constantinou, Norman Fenton and Liam Pollock** which developed a
causal Bayesian network model to determine referee bias and applied it
to the data from all matches played in the 2011-12 Premier League
season. Here is what they say about our study:
Our related work on using Bayesian networks to predict football results is discussed here.
Another recent study might just have scientifically confirmed this possible 'Ferguson Factor', entitled, 'Bayesian networks for unbiased assessment of referee bias in Association Football'. The term 'Bayesian networks', refers to a particular statistical technique deployed in this research, which mathematically analysed referee bias with respect to fouls and penalty kicks awarded during the 2011-12 English Premier League season.
The authors of the study, Anthony Constantinou, Norman Fenton and Liam Pollock found fairly strong referee bias, based on penalty kicks awarded, in favour of certain teams when playing at home.
Specifically, the two teams (Manchester City and Manchester United) who finished first and second in the league, appear to have benefited from bias that cannot be explained by other factors. For example a team may be awarded more penalties simply because it's more attacking, not just because referees are biased in its favour.
The authors from Queen Mary University of London, argue that if the home team is more in control of the ball, then, compared to opponents, it's bound to be awarded more penalties, with less yellow and red cards, compared to opponents. Greater possession leads any team being on the receiving end of more tackles. A higher proportion of these tackles are bound to be committed nearer to the opponent's goal, as greater possession also usually results in territorial advantage.
However, this study, published in the academic journal 'Psychology of Sport and Exercise', found, even allowing for these other possible factors, Manchester United with 9 penalties awarded during that season, was ranked 1st in positive referee bias, while Manchester City with 8 penalties awarded is ranked 2nd. In other words it looks like certain teams (most specifically Manchester United) benefited from referee bias in their favour during Home games, which cannot be explained by any other possible element of 'Home Advantage'.
** Constantinou, A. C., Fenton, N. E., & Pollock, L. (2014). "Bayesian networks for unbiased assessment of referee bias in Association Football". To appear in Psychology of Sport & Exercise. A pre-publication draft can be found here.What makes this result particularly interesting, the authors argue, is that for most of the season, these were the only two teams fighting for the English Premiere League title. Were referees influenced by this, and it impacted on their decision-making? Conversely the study found Arsenal, a team of similar popularity and wealth, and who finished third, benefited least of all 20 teams from referee bias at home, with respect to penalty kicks awarded. With the second largest average attendance as well as the second largest average crowd density, Arsenal were still ranked last in terms of referee bias favouring them for penalties awarded. In other words, Arsenal didn't seem to benefit much at all from the kind of referee bias that other teams were gaining from 'Home Advantage'. Psychologists might argue that temperament-wise, Sir Alex Ferguson and Arsene Wenger appear at opposite poles of the spectrum.
Our related work on using Bayesian networks to predict football results is discussed here.
Wednesday, 2 April 2014
Bayesian network approach to Drug Economics Decision Making
Consider the following problem:
A relatively cheap drug (drug A) has been used for many years to treat patients with disease X. The drug is considered quite successful since data reveals that 85% of patients using it have a ‘good outcome’ which means they survive for at least 2 years. The drug is also quite cheap, costing on average $100 for a prolonged course. The overall “financial benefit” of the drug (which assumes a ‘good outcome’ is worth $5000 and is defined as this figure minus the cost) has a mean of $4985.On seeing the data the Health Authority recommends a ban against the use of drug B. Is this a rational decision?
There is an alternative drug (drug B) that a number of specialists in disease X strongly recommend. However, the data reveals that only 65% of patients using drug B survive for at least 2 years (Fig. 1(b)). Moreover, the average cost of a prolonged course is $500. The overall “financial benefit” of the drug has a mean of just $2777.
The answer turns out to be no. The short paper here explains this using a simple Bayesian network model that you can run (by downloading the free copy of AgenaRisk)
Subscribe to:
Posts (Atom)














