New Podcast: The Hidden Algorithms Behind Modern Policing

“There’s a significant lack of transparency around police use of predictive policing systems in the UK. Most people don’t even know about their use in policing… People don’t know, if they are ever stopped and searched by police, it’s as a result of a predictive profiling or risk assessment system… So there is a significant lack of transparency, which is particularly worrying given the lack of regulation.” 

Ilyas Nagdee, Amnesty International 

Episode 12 of the Guardians of Data Podcast is out now. In this episode we discuss something that sounds like science fiction, but is already part of everyday policing in the UK; predictive policing. These are tools that use data and algorithms to help the police forecast crime, where it might happen and sometimes who might be involved. The idea is to use resources efficiently and cut crime. 

But a recent report by Amnesty International, (“Automated Racism”) argues that predictive policing tools aren’t neutral; they may be reinforcing and scaling existing inequalities. The report argues that the data they use is biased, particularly against black and racialised communities in deprived areas.  

In this conversation, we unpack what these tools actually are, how they’re being used, whether they work, and what the risks are, especially when combined with other technologies like facial recognition. 

Our guest is Ilyas Nagdee who is a human rights campaigner and the Racial Justice Director at Amnesty International UK. He has also written for The Guardian and is the co-author of a book entitled, Race to the Bottom: Reclaiming Antiracism, a critical study of anti-racist politics in the UK. 

Listen on your preferred platform via our podcast page, or download the episode directly.

This podcast is sponsored by Phaselaw – a purpose-built solution for document disclosures, like subject access requests and FOI requests. Instead of redacting PDFs one by one, or forcing litigation software to do a job it wasn’t designed for, with Phaselaw you get collection, review, and redaction in one workflow. Teams across the World are using it to cut response times from weeks to days. 

For Guardians of Data listeners, Phaselaw is offering a two-month free trial; run it on live requests, see what it does to your backlog, decide from there. No card, no commitment. 

Head to https://www.phase.law/guardians to claim your free trial.  

Previous episodes of the Guardians of Data podcast have featured Caroline Wong discussing the impact of AI on cyber security, Emma Martins explaining the importance of data protection legislation, Tahir Latif talking about responsible AI deployment and Jen Persson explaining the privacy implications of the Government’s new plans for children’s data. 

Predictive Policing and the Think Family Database

Predictive Policing, or Predictive Analytics, is increasingly being promoted as a tool to help police forces prevent crime before it happens. Supporters argue that analysing vast amounts of data can help identify vulnerable people, allocate resources more effectively and enable earlier intervention. However, a recent investigation published by WIRED magazine raises important questions about whether these systems are accurate, transparent or fair enough to justify their growing use. 

WIRED, working in partnership with the nonprofit newsroom Liberty Investigates, plus the Bristol Cable and Lighthouse Reports, obtained hundreds of pages of documentation, using FOI requests, to build a comprehensive picture of a long-running partnership between Avon and Somerset Police and Bristol City Council to develop predictive policing and safeguarding tools. It reveals how the two organisations worked together to combine public sector data, develop machine-learning models and deploy risk-scoring systems intended to support policing and child protection.  

The Think Family Database 

One of these systems was the Think Family Database which was launched in 2016. According to WIRED, the database brought together information held by multiple public bodies, creating records covering almost half a million Bristol residents.
The council played a central role by contributing and managing information from housing, education, children’s services, social care and other local authority functions, while police intelligence and crime data were also incorporated. The intention was to provide practitioners across organisations with a more complete understanding of individuals and families who might require support, enabling earlier intervention and better coordination between agencies. 

Using this shared data, the two organisations developed numerous machine-learning models designed to predict a range of outcomes. These included identifying people considered at greater risk of offending, becoming victims of crime, going missing or failing to appear in court. Other models sought to identify children who might be vulnerable to criminal or sexual exploitation. 

On paper, these objectives reflected a broader ambition shared by many public sector organisations: using data more effectively to improve services and prevent harm before it occurs. Former project leaders interviewed by WIRED argued that combining information from different public bodies could provide frontline professionals with a richer understanding of vulnerability than any single agency could achieve alone. However, the investigation suggests that the practical reality proved far more challenging.  

Accuracy and Transparency  

One of the most significant findings reported by WIRED is that at least two of the predictive models were eventually withdrawn because staff no longer trusted their results. Bristol City Council commissioned independent evaluations of the programme, and council practitioners reportedly questioned whether some of the safeguarding models were accurately identifying the children they were intended to protect. According to the investigation, staff expressed concern that some vulnerable children were no longer appearing within the highest-risk groups, while other individuals received unexpectedly high risk scores. Internal reviews ultimately concluded that the models lacked sufficient reliability to support operational decision-making, leading to their withdrawal. 

The investigation also highlights wider concerns surrounding transparency and governance. Independent reviewers reportedly found that documentation explaining how some of the predictive models had been developed was incomplete or unavailable. In some, reviewers were unable to fully assess the systems because source code, technical documentation and records describing how models had been trained or validated could not be located. 

Predictive AI systems depend heavily on the information used to train them.
If historical datasets contain gaps, inaccuracies or existing biases, those weaknesses may also be reflected in the predictions generated by the models. 
Researchers interviewed by WIRED note that some variables used within the aforementioned systems could act as indirect indicators of poverty or wider social disadvantage. Factors such as housing support, school attendance or eligibility for free school meals may correlate with vulnerability, but they may also reflect structural inequalities rather than future criminal behaviour. This raises concerns that predictive systems could unintentionally reinforce existing patterns of disadvantage rather than objectively identifying risk. 

The investigation also reports that one external audit found many of the predictive models demonstrated relatively weak performance. According to WIRED, an independent AI auditing company concluded that several models produced a high number of false positives, meaning many individuals identified as high risk would never actually experience the outcomes the models were predicting. False positives are particularly significant within policing and safeguarding because they have the potential to influence professional judgement and the allocation of limited public resources. Even where algorithmic scores do not determine decisions directly, they may shape how practitioners prioritise cases or assess individuals. 

The investigation further explores issues surrounding public awareness and consent. Many residents reportedly had little knowledge that their information had been brought together within the Think Family Database. One campaigner only discovered that his information had been included within a police offender management system after pursuing legal action to obtain details of the records held about him.  

Interestingly, former project leaders interviewed by WIRED argued that frontline professionals often relied more heavily on their own experience than on the algorithmic predictions themselves. Social workers and other practitioners reportedly viewed the models as one source of information rather than definitive guidance.
While this may have reduced the practical impact of inaccurate predictions, it also raises legitimate questions about the value of developing complex predictive systems if experienced professionals ultimately lacked confidence in the results. 

Future Use 

The investigation also highlights Bristol City Council’s role in reviewing the future of the programme. The council has since stated that the current administration no longer uses predictive analytics for policing or safeguarding decisions, with the exception of analytical work aimed at identifying young people who may become Not in Education, Employment or Training (NEET) after leaving school. The council also maintained that predictive tools never replaced professional judgement and were intended only to support practitioners rather than automate decisions. 

Automated Racism 

In the next episode of the Guardians of Data Podcast (published on Wednesday we discuss predictive policing and its impact in detail. Our guest is Ilyas Nagdee who is the Racial Justice Director at Amnesty International UK and one of the authors of Amnesty’s report into predictive policing (“Automated Racism). Ilyas helps us unpack what the predictive policing tools actually are, how they’re being used, whether they work, and what the risks are, especially when combined with other technologies like facial recognition.  

Listen to a clip here. 

Follow the podcast to be the first to know when this episode is published on Wednesday.   

Also available on Apple Podcasts, Spotify, and all major podcast platforms.

Exploring the Legal and Regulatory Challenges of AI and Chat GPT 

In our recent blog post, entitled “GDPR and AI: The Rise of the Machines”, we said that 2023 is going to be the year of Artificial Intelligence (AI). Events so far seem to suggest that advances in the technology as well legal and regulatory challenges are on the horizon.   

Generative AI, particularly large language models like ChatGPT, have captured the world’s imagination. ChatGPT registered 100 million monthly users in January alone; having only been launched in November and it set the record for the fastest growing platform since TikTok, which took nine months to hit the same usage level. In March 2023, it recorded 1.6 Billion user visits which are just mind-boggling numbers and shows how much of a technological advancement it will become. There have already been some amazing medical uses of generative AI including the ability to match drugs to patients, numerous stories of major cancer research breakthroughs as well as the ability for robots to do major surgery. 
 
However, it is important to take a step back and reflect on the risks of a technology that has made its own CEO “a bit scared” and which has caused the “Godfather of AI” to quit his job at Google. The regulatory and legal backlash against AI has already started. Recently, Italy became the first Western country to block ChatGPT. The Italian DPA highlighted privacy concerns relating to the model. Other European regulators are reported to be looking into the issue too. In April the European Data Protection Board launched a dedicated task force on ChatGPT. It said the goal is to “foster cooperation and to exchange information on possible enforcement actions conducted by data protection authorities.” Elsewhere, Canada has opened an investigation into OpenAI due to a complaint alleging the collection, use and disclosure of personal information is without consent. 

The UK Information Commissioner’s Office (ICO) has expressed its own concerns. Stephen Almond, Director of Technology and Innovation at the ICO, said in a blog post

“Data protection law still applies when the personal information that you’re processing comes from publicly accessible sources…We will act where organisations are not following the law and considering the impact on individuals.”  

Wider Concerns 

ChatGPT suffered its first major personal data breach in March.
According to a blog post by OpenAI, the breach exposed payment-related and other personal information of 1.2% of the ChatGPT Plus subscribers. But the concerns around AI and ChatGPT don’t stop at privacy law.   

An Australian mayor is considering a defamation suit against ChatGPT after it told users that he was jailed for bribery; in reality he was the whistleblower in the bribery case. Similarly it falsely accused a US law professor of sexual assault. The Guardian reported recently that ChatGPT is making up fake Guardian articles. There are concerns about copyright law too; there have been a number of songs that use AI to clone the voices of artists including Drake and The Weeknd which has since  been removed from streaming services after criticism from music publishers. There has also been a full AI-Generated Joe Rogan episode with the OpenAI CEO as well as with Donald Trump. These podcasts are definitely worth a sample, it is frankly scary how realistic they actually are. 

AI also poses a significant threat to jobs. A report by investment bank Goldman Sachs says it could replace the equivalent of 300 million full-time jobs. Our director, Ibrahim Hasan, recently gave his thoughts on this topic to BBC News Arabic. (You can watch him here. If you just want to hear Ibrahim “speak in Arabic” skip the video to 2min 48 secs!) 
 

EU Regulation 

With increasing concern about the future risks AI could pose to people’s privacy, their human rights or their safety, many experts and policy makers believe AI needs to be regulated. The European Union’s proposed legislation, the Artificial Intelligence (AI) Act, focuses primarily on strengthening rules around data quality, transparency, human oversight and accountability. It also aims to address ethical questions and implementation challenges in various sectors ranging from healthcare and education to finance and energy. 

The Act also envisages grading AI products according to how potentially harmful they might be and staggering regulation accordingly. So for example an email spam filter would be more lightly regulated than something designed to diagnose a medical condition – and some AI uses, such as social grading by governments, would be prohibited altogether. 

UK White Paper 

On 29th March 2023, the UK government published a white paper entitled “A pro-innovation approach to AI regulation.” The paper sets out a new “flexible” approach to regulating AI which is intended to build public trust and make it easier for businesses to grow and create jobs. Unlike the EU there will be no new legislation to regulate AI. In its press release, the UK government says: 

“The government will avoid heavy-handed legislation which could stifle innovation and take an adaptable approach to regulating AI. Instead of giving responsibility for AI governance to a new single regulator, the government will empower existing regulators – such as the Health and Safety Executive, Equality and Human Rights Commission and Competition and Markets Authority – to come up with tailored, context-specific approaches that suit the way AI is actually being used in their sectors.” 

The white paper outlines the following five principles that regulators are to consider facilitating the safe and innovative use of AI in their industries: 

  • Safety, Security and Robustness: applications of AI should function in a secure, safe and robust way where risks are carefully managed; 

  • Transparency and Explainability: organizations developing and deploying AI should be able to communicate when and how it is used and explain a system’s decision-making process in an appropriate level of detail that matches the risks posed by the use of the AI; 

  • Fairness: AI should be used in a way which complies with the UK’s existing laws (e.g., the UK General Data Protection Regulation), and must not discriminate against individuals or create unfair commercial outcomes; 

  • Accountability and Governance: measures are needed to ensure there is appropriate oversight of the way AI is being used and clear accountability for the outcomes; and 

  • Contestability and Redress: people need to have clear routes to dispute harmful outcomes or decisions generated by AI 

Over the next 12 months, regulators will be tasked with issuing practical guidance to organisations, as well as other tools and resources such as risk assessment templates, that set out how the above five principles should be implemented in their sectors. The government has said this could be accompanied by legislation, when parliamentary time allows, to ensure consistency among the regulators. 

Michelle Donelan MP, Secretary of State for Science, Innovation and Technology, considers that this this light-touch, principles-based approach “will enable . . . [the UK] to adapt as needed while providing industry with the clarity needed to innovate.” However, this approach does make the UK an outlier in comparison to global trends. Many other countries are developing or passing special laws to address alleged AI dangers, such as algorithmic rules imposed in China or the United States. Consumer groups and privacy advocates will also be concerned about the risks to society in the absence of detailed and unified statutory AI regulation.  

Want to know more about this rapidly developing area? Our forthcoming AI and Machine Learning workshop will explore the common challenges that this subject presents focussing on GDPR as well as other information governance and records management issues.