AI drug discovery and genomics data infrastructure in the UK: what the evidence says about results so far

Desk research on AI-designed medicines in clinical trials, the UK's petabyte-scale genomic datasets and how they are accessed, the policy and money behind them, and what it means for small biotechs and their IT suppliers.

Search the blog

Abstract

Artificial intelligence (AI) is now used at every stage of medicine discovery, and the United Kingdom holds some of the largest population genomic datasets in the world. This paper asks what the peer-reviewed and official evidence actually shows about results: how many AI-derived medicines have reached clinical trials and how they have fared, how large the UK's genomic data assets are and how they are accessed, what the government has committed in policy and money, what the data costs to store and protect, and where the claims run ahead of the evidence. The method is desk research on secondary sources: the AlphaFold papers in Nature, the published analyses of AI-native company pipelines, clinical success-rate and cost studies, the documentation of UK Biobank, Genomics England, the NHS Genomic Medicine Service and Our Future Health, government and regulator publications, and published cloud list prices, with a simple phase-probability model applied to the figures. The principal findings are that AI-discovered molecules show a Phase I success rate of 80–90% against an industry figure of 66%, but a Phase II rate of about 40% that is no better than history; that one AI-designed medicine has published randomised Phase II results and entered Phase III; that the UK's genomic assets run to tens of petabytes held in trusted research environments from which no patient-level data leaves; and that Phase II, not chemistry, remains the bottleneck the model rewards. For UK small firms the practical consequences are about secure access, not local storage.

Keywords: artificial intelligence; drug discovery; genomics; clinical trial success rates; trusted research environments; data protection; life sciences policy; cloud storage economics

1. Introduction

Two things happened in biotechnology during the past five years that were genuinely new. The first was that a computer program predicted the three-dimensional shape of proteins with "accuracy competitive with experimental structures in a majority of cases" (Jumper et al., 2021), removing a bottleneck that had lasted more than fifty years. The second was that the UK finished sequencing the whole genomes of half a million volunteers and made the result, about 30 petabytes of data, available to researchers anywhere in the world through a cloud platform in London (UK Biobank, 2023a; 2023b).

Both are information technology (IT) stories as much as biology stories. The first is a story about neural networks and graphics processors; the second about object storage, identity management and the economics of moving data. Both have been accompanied by strong claims: that AI will halve the cost of a new medicine, that genomic data will make the UK "the leading life sciences economy in Europe by 2030, and the third globally by 2035" (HM Government, 2025a). This paper asks what the evidence says so far, as distinct from what is hoped.

The question matters to more than pharmaceutical companies. The UK life sciences industry had a turnover of £146.9 billion and employed 359,600 people in 2023/24, and 94% of its companies were small or medium-sized enterprises (SMEs) (Department for Science, Innovation and Technology and Office for Life Sciences, 2025). Most of those firms do not own a supercomputer or a sequencer. They buy IT, and the way the sector's data is now held changes what they should buy. The contribution of this paper is to put the peer-reviewed results, the official documentation of the UK's data assets and the published prices side by side, and to draw out what follows for a small biotech, a contract research organisation (CRO) or an IT supplier to either. Section 2 reviews the literature; Section 3 sets out the method and a simple pipeline model; Section 4 gives the findings; Section 5 discusses them, including the strongest counter-argument and a section for UK SMEs; Sections 6 and 7 give limitations and conclusions.

2. Literature review

2.1 Structure prediction: the clearest scientific result

The AlphaFold papers are the anchor of the AI-in-biology literature because their claims were tested blind. Jumper et al. (2021) note that the experimental structures of "around 100,000 unique proteins" had been determined, "a small fraction of the billions of known protein sequences", and that determining a single structure took "months to years of painstaking effort". Their network, validated in the 14th Critical Assessment of protein Structure Prediction (CASP14), "can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known". Abramson et al. (2024) extended the method to complexes of proteins with small molecules, nucleic acids and antibodies, reporting "substantially improved accuracy over many previous specialized tools", including far greater accuracy than docking tools on the PoseBusters protein–ligand benchmark. Structure prediction is one step in drug discovery, not the whole of it, but it is the step where the evidence is least in dispute.

2.2 AI-native pipelines: from counting assets to counting successes

The Boston Consulting Group analyses by Jayatunga and colleagues are the most cited attempts to measure output rather than promise. Jayatunga et al. (2022) examined 24 "AI-native" drug discovery companies and found, for the 20 with disclosed pipelines, "~160 disclosed discovery programmes and preclinical assets and about 15 assets in clinical development", with pipelines growing at "around 36%" a year; against that, the top 20 pharmaceutical companies' in-house pipelines held about 330 discovery and preclinical assets and about 430 assets in Phase I (Boston Consulting Group, 2022). The AI companies' combined pipeline was, on their estimate, "equivalent to 50% of the in-house discovery and preclinical output" of big pharma, and several programmes had gone from discovery to the clinic in under four years.

Two years later the same group asked the harder question. Jayatunga et al. (2024) analysed the clinical pipelines of AI-native biotechs and found that "In Phase I we find AI-discovered molecules have an 80-90% success rate, substantially higher than historic industry averages", which they read as evidence that "AI is highly capable of designing or identifying molecules with drug-like properties". In Phase II, however, "the success rate is ∼40%, albeit on a limited sample size, comparable to historic industry averages". That asymmetry, chemistry improved and biology unchanged, is the central empirical fact of this paper.

2.3 What "historic industry averages" are

Wong, Siah and Lo (2019) provide the largest independent estimate, drawn from 406,038 clinical-trial records covering more than 21,143 compounds between 2000 and 2015. They find that "13.8% of all drug development programs eventually lead to approval", with phase transition probabilities of 66.4% from Phase I to II, 48.6% from Phase II to III and 59.0% from Phase III to approval, a 3.4% overall success rate in oncology and 20.9% outside it, and higher success where trials used biomarkers to select patients. DiMasi, Grabowski and Hansen (2016), using confidential data from 10 companies on 106 drugs, estimate the out-of-pocket cost per approved compound at $1,395 million in 2013 dollars and the capitalised pre-approval cost at $2,558 million at a 10.5% real discount rate, rising to $2,870 million with post-approval work, with costs growing "at an annual rate of 8.5% above general price inflation". The two studies disagree on method and sample but agree that most of the money is spent on candidates that fail, which is exactly the lever AI is claimed to move.

2.4 Genomic data as an infrastructure problem

Stephens et al. (2015) argued that genomics would become "on par with or the most demanding" of the big-data domains in acquisition, storage, distribution and analysis, projecting that "between 100 million and as many as 2 billion human genomes could be sequenced by 2025" needing "2–40 exabytes of storage capacity", with sequence output "doubling approximately every seven months" and about 100 gigabases of raw data collected for every 3 billion bases of finished genome. The UK's response has been to concentrate the data in a few custodians and to bring researchers to the data rather than send data to researchers. UK Biobank describes its Research Analysis Platform (RAP) as "a restricted, sophisticated and user-friendly cloud-based tool" (UK Biobank, n.d. b); Genomics England states that "No individual patient-level data can be exported from the Research Environment, instead only results of analyses can be exported" (Genomics England, n.d. a). This is the trusted research environment (TRE) model, and it is the IT architecture around which UK policy is now built.

2.5 Regulation and data protection

The Information Commissioner's Office (ICO) treats genomic data as special category data that "is difficult to effectively anonymise, given the uniqueness of the information", and notes that disclosure "also affects related family members" (ICO, 2024a). Its executive director for regulatory risk put it more directly: "Genomic information is arguably the most sensitive and revealing information a person has, with major implications for not only individuals but their families" (ICO, 2024b). The National Cyber Security Centre (NCSC) identifies three adversaries for research organisations, "State actors looking to steal technology", "Competitors seeking commercial advantage" and "Criminals looking to profit from weak security" (NCSC, 2025a). On the AI side, the Medicines and Healthcare products Regulatory Agency (MHRA) launched the AI Airlock, a regulatory sandbox for AI as a medical device, in May 2024 (MHRA, 2024). It is worth being precise that the Airlock concerns AI devices used in care, not AI used to design medicines; an AI-discovered drug goes through the ordinary clinical route.

3. Method

This is desk research on secondary data. No interviews, proprietary databases or unpublished figures were used. Sources were ranked in a fixed order: peer-reviewed journals; UK government, regulator and NHS publications; the custodians' own documentation (UK Biobank, Genomics England, Our Future Health); company announcements, used only for facts about the company itself; and vendors' published list prices. Every number in the paper is taken from a document retrieved for this study and listed in the references.

Two simple models are applied. The first is the standard multiplicative pipeline model:

P=pI·pII·pIII, nk=1j=kIIIpj
(1)

Here P is the probability that a candidate entering Phase I is eventually approved; pI, pII and pIII are the probabilities of passing Phases I, II and III; and nk is the expected number of candidates that must enter phase k to yield one approval, obtained by dividing one by the product of the transition probabilities from phase k onwards. So nI = 1/P.

The second links that to cost and to storage:

E=k=IIIInkck, S=12·g·b·r
(2)

Here E is the expected development cost per approved medicine, ck is the cost of taking one candidate through phase k, and nk is as above; the sum runs over the three clinical phases. In the storage term, S is the annual cost of holding a dataset, g is the number of genomes, b is the stored size per genome in gigabytes, r is the vendor's price per gigabyte per month, and 12 converts to a year. Two assumptions are stated up front. Equation 1 treats phases as independent, which slightly overstates P: multiplying Wong, Siah and Lo's (2019) phase rates gives 19%, whereas their path-by-path estimate is 13.8%, the difference arising because candidates that fail are not a random sample of those that enter each phase.8%); the model is used here for comparison between scenarios, not for point forecasts. Equation 2 excludes retrieval, request, egress and compute charges, all of which are real, and treats a petabyte as one million gigabytes.

4. Findings

4.1 AI-discovered medicines in the clinic

Figure 1 puts the two Jayatunga et al. (2024) findings against the Wong, Siah and Lo (2019) baselines.

Clinical phase transition probabilities: industry 2000 to 2015 versus AI-discovered molecules Grouped bar chart. Industry (Wong, Siah and Lo 2019): Phase I to II 66.4 percent, Phase II to III 48.6 percent, Phase III to approval 59.0 percent, overall Phase I to approval 13.8 percent. AI-discovered molecules (Jayatunga et al. 2024): Phase I 80 to 90 percent shown as 85 with a range bar, Phase II about 40 percent, no data for Phase III or overall. Probability of passing each clinical phase (%) 0 20 40 60 80 100 Success probability (%) 66.4 80–90 Phase I to II 48.6 ≈40 Phase II to III 59.0 no data Phase III to approval 13.8 no data Phase I to approval Industry, 2000–2015 AI-discovered molecules Source: Wong, Siah and Lo (2019), Table 1; Jayatunga et al. (2024). AI Phase I plotted at 85 with the reported 80–90 range; AI Phase II reported as about 40 on a small sample.
Figure 1. Figure 1. Probability of passing each clinical phase: industry 2000–2015 versus AI-discovered molecules. Sources: Wong, Siah and Lo (2019); Jayatunga et al. (2024).

The Phase I result is striking and the Phase II result is not. Applying Equation 1 with the industry figures gives P = 0.664 × 0.486 × 0.590 = 19.0% and nI = 5.3 Phase I starts per approval (7.2 on Wong's own 13.8% figure). Replacing the Phase I rate with 85%, the midpoint of the AI range, and leaving the rest unchanged gives P = 24.4% and nI = 4.1: a 28% improvement in the odds, but only in the odds of surviving the cheapest phase. Table 1 works through the cases.

Table 1. Pipeline scenarios under Equation 1. Baseline transition rates from Wong, Siah and Lo (2019); AI rates from Jayatunga et al. (2024); the 60% Phase II case is hypothetical.
ScenariopIpIIpIIIP (Phase I to approval)Phase I starts per approvalPhase II starts per approval
Industry baseline (product of rates)66.4%48.6%59.0%19.0%5.33.5
AI Phase I only (85%)85%48.6%59.0%24.4%4.13.5
AI Phase I and reported Phase II (85%, 40%)85%40%59.0%20.1%5.04.2
Hypothetical: AI improves Phase II to 60%85%60%59.0%30.1%3.32.8

The third row is the sobering one. If the AI Phase II rate of about 40% holds up on a larger sample, the number of Phase I starts per approval falls only from 5.3 to 5.0, and the number of Phase II starts, the expensive ones, rises. Equation 2 makes the point in money: nII and nIII are unchanged by anything that happens in Phase I, so the saving from better chemistry is bounded by the Phase I share of the $2,558 million capitalised cost per approval that DiMasi, Grabowski and Hansen (2016) estimate. That study does not publish a phase split, so this paper does not put a number on the bound; the qualitative conclusion, that the prize lies in Phase II, stands without it.

Company disclosures fill in the picture at the level of individual medicines, and are used here only for facts about those companies. The most complete published result is rentosertib, a TNIK inhibitor for idiopathic pulmonary fibrosis for which, the authors state, both the target and the compound were identified with generative AI. In a randomised, double-blind, placebo-controlled Phase IIa trial of 71 patients over 12 weeks, the 60 mg once-daily group showed a mean change in forced vital capacity of +98.4 ml (95% confidence interval 10.9 to 185.9) against −20.3 ml on placebo, with similar adverse-event rates across arms (Xu et al., 2025). A Phase III trial of 320 participants across 47 centres in China began in July 2026; the sponsor reports 31 preclinical candidate nominations, 13 investigational new drug clearances and 8 ongoing Phase I trials (Insilico Medicine, 2026). Recursion reported five clinical programmes and cash of $753.9 million at 31 December 2025, with its MEK1/2 inhibitor REC-4881 producing a 43% median reduction in total polyp burden after 12 weeks (Recursion Pharmaceuticals, 2026). Isomorphic Labs raised $600 million in March 2025 to "advance therapeutic programs into the clinic" (Isomorphic Labs, 2025); the government's sector plan names it as its example of UK AI drug discovery (HM Government, 2025a). At the time of writing, then, the evidence base is one positive randomised Phase II result now in Phase III, a handful of early-phase signals, and a great deal of well-funded preclinical work. No AI-discovered medicine has been approved.

4.2 The UK's genomic data assets

Figure 2 lays out the four principal assets and the access model they share.

UK genomic data assets and the trusted research environment access model Schematic. Four data assets (UK Biobank, Genomics England, NHS Genomic Medicine Service, Our Future Health) feed a trusted research environment in which data is analysed but never leaves; outputs pass through an airlock review before results are released to the researcher. UK genomic data assets and how they are accessed UK Biobank Genomics England NHS Genomic Medicine Our Future Health 500,000 participants whole genomes + exomes about 30 PB of sequence 22,000+ researchers cloud, London region 140,000+ whole genomes rare disease and cancer 50 PB from 100,000 Genomes target 500,000 by 2030 linked NHS records Service (NHS England) 7 laboratory hubs national test directory 40,968 genomes in research release v5 (Aug 2025) up to 5 million people 1.93m questionnaires (2026) blood sample + measures £354m government backing de-identified, secure access Trusted research environment (secure cloud or data centre in the UK) Data is analysed inside and never downloaded; the researcher pays an access fee plus compute and storage charges. Identity checks, approved project, audit logging; no copy-and-paste out; no individual-level export. Airlock review by the data custodian Every file leaving is checked against consent, ethics and data protection rules before release. What leaves: results, not records Summary statistics, models and code return to the researcher; patient-level data stays behind. Source: UK Biobank (2023a, 2023b, n.d. b, n.d. c); Genomics England (n.d. a, 2025, 2026); Amazon Web Services (n.d.); NHS England (n.d. b); Our Future Health (n.d., 2026); HM Government (2025a).
Figure 2. Figure 2. UK genomic data assets at scale and the trusted research environment access model. Sources: UK Biobank (2023a, 2023b, n.d. b, n.d. c); Genomics England (n.d. a, 2025, 2026); Amazon Web Services (n.d.); NHS England (n.d. b); Our Future Health (n.d., 2026); HM Government (2025a).

UK Biobank holds whole genome and exome sequences for 500,000 people, imaging on 100,000 and proteomics on 54,000, and reports over 22,000 researchers in more than 60 countries and more than 18,000 peer-reviewed papers (UK Biobank, n.d. a; n.d. b). The whole-genome programme, sequencing 491,554 participants, took from September 2019 to early 2022 and was funded by a £200 million consortium of Amgen, AstraZeneca, GSK, Johnson & Johnson, government and Wellcome (UK Biobank, 2023a). The release amounted to "about 30 petabytes of data, with nearly 18 million individual-level and 600,000 population-level files" on a platform hosted on Amazon Web Services in the London region and operated with DNAnexus (UK Biobank, 2023b; 2023a). Access is tiered: £3,000, £6,000 or £9,000 for three years depending on data type, £500 for students and researchers from lower-income countries, with compute, uploaded storage and data egress charged separately (UK Biobank, n.d. c). The platform is, at the time of writing, closed "whilst we implement necessary changes", with phased re-opening intended from September 2026 (UK Biobank, n.d. b); a reminder that a single national platform is also a single point of dependency.

Genomics England passed the 100,000-genome mark of the 100,000 Genomes Project in 2018, completing recruitment in December that year, from around 85,000 NHS patients with rare disease or cancer; in the rare-disease pilot cohort of over 4,000 people from over 2,000 families, whole genome sequencing gave a new diagnosis to 25% (Genomics England, n.d. b). Its National Genomic Research Library now holds "140,000+ whole genomes" with linked hospital, cancer-treatment and mental-health records (Genomics England, n.d. a), and the sector plan commits over £650 million to reach "over 500,000 genomes" by 2030 (HM Government, 2025a). The 100,000 Genomes Project alone produced 50 petabytes, now being migrated to the public cloud (Amazon Web Services, n.d.). Genomics England reports that the average 100,000 Genomes participant has contributed to 137 research projects, and that every output passes through an Airlock review before it leaves the environment (Genomics England, 2026).

The NHS Genomic Medicine Service delivers testing through seven Genomic Laboratory Hubs under a single national test directory, with the stated aim of being "the first national health care system to offer whole genome sequencing as part of routine care" (NHS England, n.d. a; n.d. b). Its fifth research data release, in August 2025, contained 40,968 genomes from 38,728 participants, 36,500 of them rare-disease cases and 4,468 paired tumour and germline cancer genomes (Genomics England, 2025). This is the point where research data and care data meet, and where the data-protection questions in Section 4.4 become live.

Our Future Health is recruiting "up to five million people" who give a questionnaire, a blood sample and physical measurements, with researchers applying "to study the de-identified information in a highly secure online data storage system" (Our Future Health, n.d.); a July 2026 analysis used questionnaire data from 1,929,744 participants (Our Future Health, 2026), and the sector plan backs the programme with up to £354 million (HM Government, 2025a).

Alongside the research assets sits the operational NHS Federated Data Platform, supplied by a consortium led by Palantir Technologies UK under a contract that "formally began in March 2024", with three years committed of a possible seven; the supplier cannot "commercialise or market NHS data, even on an anonymised basis" (NHS England, n.d. c). The government and Wellcome are also funding, at up to £600 million, a Health Data Research Service at the Wellcome Genome Campus, alongside a target to cut clinical-trial set-up from over 250 days in 2022 to 150 days by March 2026 (Prime Minister's Office, 2025).

4.3 What the data costs to keep

Figure 3 applies the storage term of Equation 2 to a dataset the size of UK Biobank's genome release, using published list prices.

Annual cloud storage cost of a 30 petabyte genomic dataset by storage class Bar chart. At Google Cloud list prices for the London region, storing 30 petabytes for a year costs about 7.2 million US dollars in Standard storage, 3.6 million in Nearline, 1.44 million in Coldline and 0.43 million in Archive. Cost of storing 30 PB of genomes for one year, by storage class (US$ million) 0 2 4 6 8 US$ million per year 7.2 Standard 3.6 Nearline 1.44 Coldline 0.43 Archive Per GB-month: $0.020, $0.010, $0.004, $0.0012 Source: Google Cloud (2026) published list prices (default region), hourly rate × 730 h; size from UK Biobank (2023b). Storage only: excludes retrieval, requests, egress and compute. 30 PB treated as 30 million GB.
Figure 3. Figure 3. Annual list-price cost of storing 30 petabytes of genomic data by storage class, Google Cloud published default-region prices. Sources: Google Cloud (2026); UK Biobank (2023b).

Google Cloud's published list prices (its default-region display; the London region is priced separately and may differ) put Standard storage at $0.000027397 per gibibyte-hour, Nearline at $0.000013699, Coldline at $0.000005479 and Archive at $0.000001644, and data transfer to the internet at $0.12 per gibibyte for the first 10 tebibytes a month (Google Cloud, 2026). Multiplied by 730 hours, those are about $0.020, $0.010, $0.004 and $0.0012 per gigabyte-month. For 30 petabytes, Equation 2 gives roughly $7.2 million a year in Standard storage and $0.43 million in Archive. Moving the whole dataset out once, at the headline egress rate, would cost about $3.6 million, which is the economic reason the data does not move and the researcher does. The same arithmetic scales down: 30 petabytes across 500,000 genomes is about 60 gigabytes per genome, so a small firm holding 1,000 of its own genomes (60 terabytes) would pay about $1,200 a month in Standard storage or about $72 in Archive, before compute and before the backups that any sensible firm keeps outside its primary provider. Those figures are list prices; committed-use and volume discounts apply at scale and are not published in a form this paper can cite.

4.4 Data protection and security

Three features of genomic data change the IT design. It cannot be reliably anonymised, it identifies relatives who never consented, and it does not change over a lifetime, so a breach has consequences that outlast any password reset (ICO, 2024a; 2024b). The custodians' response is architectural: keep the data in one place, bring the analysis to it, log everything, and review outputs before release (Genomics England, n.d. a; UK Biobank, n.d. b). The threat is not hypothetical. The NCSC handled 429 incidents in the year to August 2025, of which 204 were nationally significant, up from 89 the year before, and 18 highly significant, up from 12 (NCSC, 2025b). Its guidance for research organisations and start-ups, under the Trusted Research and Secure Innovation banners, is written for small and medium-sized organisations and the public sector as well as security professionals (NCSC, 2025a).

4.5 Policy and money

The Life Sciences Sector Plan of July 2025 commits "over £2 billion" across the spending review period: up to £600 million for the Health Data Research Service, up to £520 million for the Life Sciences Innovative Manufacturing Fund, over £650 million for Genomics England, up to £354 million for Our Future Health and up to £20 million for UK Biobank's protein-measurement work, with the MHRA funded to expand the AI Airlock (HM Government, 2025a; 2025b). The Airlock itself has completed two phases with 11 innovators across 7 regulatory challenges and received multi-year funding in April 2026 (MHRA, 2026). The Office for Life Sciences, which sits across the Department of Health and Social Care and the Department for Science, Innovation and Technology, is the coordinating body (Office for Life Sciences, n.d.).

Private money is more cautious. UK biotech raised £1.79 billion of venture capital in 58 deals in 2025, down 13.2% on 2024, although the average deal rose from £18.7 million to £30.8 million and the UK took 30% of European venture financing; there were no initial public offerings, in "the fifth year of restricted activity" for the listing window (BioIndustry Association, 2026). The BioIndustry Association describes the Isomorphic Labs and Verdiva Bio rounds as the landmark investments of the year, together accounting for almost 47% of the capital raised in 2025 (BioIndustry Association, 2026). The sector statistics show why the money is concentrated: biopharmaceutical firms are 40% of companies but 67% of turnover, and SMEs, 94% of companies, generate 32% of turnover (Department for Science, Innovation and Technology and Office for Life Sciences, 2025).

5. Discussion

5.1 What the evidence supports

Three conclusions are supported. First, AI has demonstrably improved the front end of discovery: protein structures that took years now take minutes (Jumper et al., 2021; Abramson et al., 2024), AI-native pipelines are large and growing (Jayatunga et al., 2022), and molecules from them survive Phase I at 80–90% rather than 66% (Jayatunga et al., 2024; Wong, Siah and Lo, 2019). Second, the UK has built genuinely large data assets and a coherent way of accessing them; 30 and 50 petabytes are not marketing figures, and the trusted research environment with an airlock is now the common pattern across UK Biobank, Genomics England and Our Future Health. Third, the government has put specific sums behind specific institutions, which is more than most sector plans do.

5.2 What it does not yet support

The evidence does not support the claim that AI has changed the economics of drug development. Under Equation 1 a Phase I improvement alone raises the odds of approval by about a quarter, and under Equation 2 the cost saving is confined to the Phase I share of spending. The reported Phase II rate of about 40% is, in the authors' own words, "comparable to historic industry averages" (Jayatunga et al., 2024); Phase II is where a medicine meets human biology, and there is no published evidence yet that AI-selected targets fail there less often. One randomised Phase IIa trial with a positive lung-function signal in 71 patients (Xu et al., 2025) is a real result and a small one. The literature's own language is careful: "early signs of the clinical potential" (Jayatunga et al., 2024), "a coming wave?" with a question mark (Jayatunga et al., 2022). Public discussion has tended to drop the question mark.

Nor does the evidence show that data scale alone produces medicines. UK Biobank's 18,000 papers are a scientific return, not a commercial one, and the route from a genetic association in 500,000 volunteers to an approved drug still runs through the same three phases. Wong, Siah and Lo (2019) do offer one encouraging finding: trials that used biomarkers to select patients had higher overall success probabilities than trials that did not. Genomic data is the raw material of such biomarkers, and that is the mechanism by which the UK's assets are most likely to move the numbers in Table 1: by raising pII, which is the lever that matters.

5.3 The strongest counter-argument

The best case against this paper's caution runs as follows. Success rates lag by construction: a molecule designed in 2021 cannot have a Phase III result in 2026, so a Phase II rate measured today reflects the earliest and crudest AI methods, and the 80–90% Phase I rate already shows the methods work where they have had time to be tested. Rentosertib, target and molecule both AI-derived, went from discovery to a randomised Phase IIa result and into Phase III inside the four-year window Jayatunga et al. (2022) described. The largest structure-prediction advance in history is only two years into being applied to complexes and antibodies (Abramson et al., 2024). Isomorphic's $600 million, and the government's £2 billion, are bets by people who have read the same numbers. On this view the correct reading of Figure 1 is not "Phase II unchanged" but "Phase II not yet measured".

The reply is that the counter-argument is a forecast and this paper is an audit. Both can be right. The audit finds one approved-track result and a Phase II rate no better than history; the forecast says wait. What the audit rules out is the present-tense claim, common in prospectuses and policy papers, that AI has already made medicines cheaper or faster to approve. It has made them faster to find.

5.4 Implications for UK small and medium businesses

For a lab-scale biotech, a CRO or an IT supplier to either, the findings translate into four practical points.

Storage is not the problem; access is. The national datasets are used inside trusted research environments and cannot be downloaded. A ten-person firm does not need petabytes; it needs verified identities for its named researchers, the fee budget (£3,000 to £9,000 for three years at UK Biobank, plus metered compute), a way to get its own code into the environment, and an airlock-compatible habit of taking out summaries rather than rows. The IT supplier's job is the endpoint the researcher connects from, not the server the data sits on.

Its own data is another matter. A firm's own sequencing, assay and trial data is special category data on the ICO's reading, identifies relatives, and is of interest to the three adversaries the NCSC names. The 3-2-1 backup rule applies with the addition that one copy must be encrypted at rest under keys the firm controls, and that egress charges make "we will just download it all if we need to" an expensive assumption; the worked example in Section 4.3 is the place to start a budget. Our cloud services page describes the portable, backed-up pattern we recommend for exactly this reason.

Certification opens doors. A supplier to a sector whose custodians run identity checks and audit logging on every user should expect to be asked how it secures its own; the April 2026 Cyber Essentials update is the sensible minimum. The NCSC's Secure Innovation material offers a free, short action plan for start-ups (NCSC, 2025a).

Read the pipeline, not the press release. A CRO or supplier choosing which AI-native clients to extend credit to should ask where each programme sits in Table 1. Cash on hand and Phase II data are the numbers that matter; Phase I success and preclinical nominations are, on the evidence here, close to the industry norm for well-run chemistry.

6. Limitations

This is a desk study of published documents. The AI success rates rest on one research group's analyses of company-disclosed pipelines, with a Phase II sample the authors call limited; disclosed pipelines omit quiet failures, which biases rates upward. Equation 1 assumes independent phases and is used for comparison, not prediction. Equation 2 uses one vendor's list prices, ignores compute, retrieval, request and egress charges except where stated, and treats a petabyte as one million gigabytes. Company figures (Insilico, Recursion, Isomorphic) are self-reported and not independently audited here. The sector statistics for 2023/24 are, by the publisher's own note, not comparable with earlier years because of a change in method. Several custodian pages are undated and may change; the UK Biobank platform status was as published at the time of retrieval. Two sources were not retrievable in primary form: the DiMasi, Grabowski and Hansen (2016) phase-level cost split, so the Phase I cost bound is left unquantified, and the UK Biobank imaging figure, which its own pages give as both 90,000 and 100,000.

7. Conclusion

AI has changed how quickly a candidate medicine can be found and how often it survives its first human trial; the peer-reviewed record shows Phase I success of 80–90% for AI-discovered molecules against an industry rate of 66%. It has not yet changed the odds in Phase II, where the reported rate of about 40% matches history, and so under a simple pipeline model it has improved the probability of approval by roughly a quarter and the cost by at most the Phase I share. One AI-designed medicine has a published randomised Phase II result and a Phase III trial under way; none is approved. The UK, meanwhile, holds tens of petabytes of population genomic data in trusted research environments with a common access model, funded by more than £2 billion of public commitments and £1.79 billion of 2025 venture capital, in a sector where 94% of firms are small. For those firms the practical lesson is that the data will not come to them: their investment belongs in secure access, certified endpoints, encrypted and portable custody of their own data, and a clear-eyed reading of which of their partners has crossed Phase II.

References

How to cite this paper

suj (2026) AI drug discovery and genomics data infrastructure in the UK: what the evidence says about results so far. IT Support World blog, 25 August 2026. Available at: https://itsupportworld.co.uk/blog/information-technology-biotechnology-ai-drug-discovery-genomics-uk.html (accessed [date]).

Written by suj. Facts checked against the cited sources at time of publication; regulations, prices and dates change — ask us if you need the current position.

© 2026 suj and IT Support World Limited. All rights reserved. Short quotations with attribution and a link are welcome; reproduction of the whole or substantial parts requires written permission from info@itsupportworld.co.uk.

Want this handled for your business?

A short conversation with an engineer — not a salesperson — is the fastest way to find out what you actually need.

Vision House, 3 Dee Road, Richmond TW9 2JN Registered UK company no. 09064078 No cookies, no trackers on this site