Francisco Javier Campos

selected builds

Work that can be opened, run, and challenged.

These projects connect AI research with delivery: concrete problems, inspectable implementation, automated tests, and explicit limitations. They are working artefacts, not claims of production readiness.

01 / safety evaluation

RSI Eval Lab

A runnable research prototype for auditing recursively self-improving AI lineages when the system may have incentives to change its evaluator or weaken oversight.

Contribution
Design, implementation, tests, and research agenda
Evidence
Deterministic auditor, synthetic traces, CI, optional Inspect task
Limits
Research harness; not containment or a deployment safety case
rsi-eval
$ rsi-eval examples/safe_run.json

Run: safe-adaptive-lineage
Verdict: PASS
Held-out delta: +0.300
Efficiency: +0.0800 / 1k tokens
Next challenge level: 3
Findings: none
evaluator hashshutdowncontainmentaudit trail
02 / agent research

AgentMisalignment

A propensity benchmark for misaligned behaviour in LLM-based agents, developed with the Cambridge AI Safety Hub and contributed to the UK AI Security Institute's Inspect Evals repository.

Contribution
Benchmark co-development, experiments, analysis, and paper
Evidence
Realistic agent environments across five behavioural categories
Finding
System-prompted persona can shift behaviour as much as model choice
01Shutdown resistance
02Oversight evasion
03Sandbagging
04Deception
05Power-seeking
03 / engineering transformation

AI Engineering Journey

A dependency-free assessment that turns seven organisational capabilities into a transparent maturity score and phased transformation roadmap.

Contribution
Operating model, scoring method, CLI, example, and tests
Evidence
Validated input contract and deterministic roadmap generation
Limits
A diagnostic for prioritisation, not a maturity certification
product4
workflow3
evaluation2
architecture3
platform4
governance2
learning3
04 / board governance

Board AI Governance Toolkit

A lightweight toolkit that turns board-level AI oversight into explicit decisions, accountable owners, evidence, and review dates.

Contribution
Risk contract, validator, dashboard generator, and templates
Evidence
Deterministic scoring, overdue-review checks, and automated tests
Limits
Must be adapted to legal duties, risk appetite, and jurisdiction
riskownerlevelreview
AI-002CPOhigh15 Oct
AI-001Opsmediumoverdue
AI-003CTOmedium30 Nov

1 decision required · 1 overdue review

05 / privacy-conscious product

Garmin Running Dashboard

An end-to-end Streamlit product that parses Garmin exports into training analysis while keeping personal health and location data out of the public repository.

Contribution
Product design, ETL, analysis, UI, synthetic demo, and tests
Evidence
156-run GPS-free demo, desktop/mobile verification, and CI
Privacy
Uploads processed in memory; local exports and outputs ignored
Garmin Running Dashboard showing synthetic weekly training data
Synthetic data only - no personal location or health records.

working contribution

What I bring to a focused build

01

Turn an ambiguous institutional problem into a testable service hypothesis.

02

Build and integrate an agentic prototype with explicit data and permission boundaries.

03

Design task-specific evaluations, failure tests, and human escalation.

04

Connect technical choices to ownership, risk, and measurable outcomes.

05

Present the result clearly to technical, policy, operational, and executive stakeholders.

contact

For executive technology mandates, public-interest AI collaboration, board advisory, or the safe delivery of agentic systems - start a conversation.