Safety AI research map
Public research links across agent misalignment, loss of control, safety cases, Assurance 2.0, and externally reviewed evidence.
Francisco Javier Campos Zabala
Group CTO and Chief AI Officer · AI safety researcher · hands-on builder
public-interest ai + accountable delivery
I turn complex public and enterprise problems into working prototypes, delivery systems, and controls - combining executive operating experience with current AI engineering and safety research.
25+ years leading technology, data, and AI organisations. Group CTO & Chief AI Officer at Cape.io. Author of Autonomous Minds. Latest safety-case paper at ICML 2026 and co-author of AgentMisalignment (ICLR 2026).
Explore AI safety researchpublic-interest delivery
Useful AI starts with the service outcome and ends with an operating system that people can trust, challenge, and improve.
Identify the user, operational constraint, baseline, and decision that needs to improve.
Use agents where permissions, data boundaries, and human accountability are explicit.
Test quality, failure, misuse, equity, privacy, and recovery under realistic conditions.
Define ownership, workflow, controls, monitoring, and a learning loop for operations.
playbooks
Public research links across agent misalignment, loss of control, safety cases, Assurance 2.0, and externally reviewed evidence.
A practical journey from coding assistance to agentic engineering: specs, harnesses, eval loops, cost constraints, and evidence-based review.
current evidence
Selected as one of 20 expert members while at Experian.
Contributed to the Forum's February 2022 final report, which mapped the barriers, risks, and potential mitigations for adopting AI in financial services across three connected areas: data, model risk, and governance. It also set out examples of good practice to support safe adoption.
Applying structured claims, arguments, and evidence to challenge a frontier model safety case and make it usable for real decisions.
Watch the ICML presentationOpen work on shutdown resistance, oversight evasion, evaluator integrity, containment, and the limits of fixed benchmarks.
Review the evaluation workTeaching enterprise AI strategy, agentic organisations, and safety governance to senior leaders as an MIT xPRO guest industry expert.
Teaching material remains privateA practical engineering journey through specification, delegation, evaluation, review, cost, and operational control.
Explore the delivery journeyselected work
External review of DeepMind's public scheming-inability safety case for Gemini 2.5, applying the Assurance 2.0 / Claims-Argument-Evidence methodology. Translates the review into recommendations for AI developers, AISIs and regulators on how frontier safety cases should be structured to support external scrutiny.
A propensity benchmark covering shutdown resistance, oversight evasion, sandbagging and power-seeking on frontier models. Key finding: persona system prompts shift misalignment more than the choice of model itself, with direct implications for how enterprises configure deployed agents.
How agentic AI predicts and learns to enable productivity and empowerment. Governance patterns and human-machine collaboration models for the next decade.
A practical blueprint for executives scaling AI responsibly across the enterprise — from strategy to production deployment.
live
speaking
about
Francisco Javier Campos Zabala has spent 25+ years building and scaling AI, data, product, and platform organisations across AdTech, FinTech, SaaS, and regulated environments - at Group CTO and CIO level inside Cape.io, Fenestra, Experian, Kantar, GroupM, and Havas. At Cape.io he is designing agentic operating models across engineering and business functions. His research with the Cambridge AI Safety Hub and Arcadia Impact covers agent misalignment and external review of frontier safety cases. While at Experian, he was one of 20 expert members of the Bank of England and FCA AI Public-Private Forum, contributing to its final report on safe AI adoption across data, model risk, and governance. He has also contributed to industry work on the EU AI Act. He is the author of Autonomous Minds and Grow Your Business with AI.
writing
I told my engineering team to spawn parallel subagents. I left out the cost, which was printed in the same document. The document was Anthropic’s post on building multi-agent systems. I carried across the exciting line: parallel agents can cover more ground than a single agent working within its context limits. I did not carry...
Everyone mocked Apple for losing the AI race. It may be the only company whose AI position survives the next two years. From a first-principles perspective, three things stand out: no frontier model moat has survived for long, unified memory changes the economics of local AI, and open models are narrowing the gap to the...
I gave Claude Fable 5 a full week as my daily driver. By Friday I was pricing the switch to OpenAI, and the refusal that pushed me there had nothing to do with safety. From a first-principles perspective, alignment is a control mechanism, a steering wheel. I got a locked door. I asked it to...
The era of “pick the best model and win” is officially over. I watched a room of global AI leaders bury it in London. At AI World Congress, every hallway conversation circled the same conclusion: the model was never the moat. The advantage sits in what you build around it. Everyone rents the same frontier...
The era of the bigger model is officially dead. The new battleground is the harness. A coding agent is a stochastic generator wrapped in a verifier. The model proposes. The harness disposes. The harness is only as honest as the eval set behind it. A model output is a hypothesis. Without a ground-truth check -...
What survives when AI starts improving itself? That was the sharpest question at MIT xPRO’s AI for Senior Executives cohort, where the room was past the “should we adopt AI” debate. They were asking what compounds for them specifically once the model layer starts compounding for everyone. Three layers sit above the model: proprietary data,...
contact
For executive technology mandates, public-interest AI collaboration, board advisory, or the safe delivery of agentic systems - start a conversation.