Oxford Edge and G-Research have come together to curate a set of ambitious open problems at the frontier of AI and machine learning, from building more capable agents and LLMs to advancing AI for science, strengthening the infrastructure it runs on, and making increasingly powerful systems safer, more reliable, and easier to trust.

If you want to get stuck into a genuinely hard problem, build something ambitious with a team, and see how far you can push it, we’re looking for you.

Who can apply and how

AI/ML Challenges is open to any matriculated student at Oxford, whatever your degree or subject. You don’t need to arrive as an expert: we’re looking for curiosity, technical ambition, and a willingness to tackle difficult problems.

Places are limited to 30.

Applications close 5 PM on 29 October.

Apply here.
Students Working

How it works

The programme kicks off in the first week of November (Michaelmas Week 4). Successful applicants will be introduced to the challenges and form teams around the problems they find most compelling.

At kick-off, you’ll get a detailed brief for each challenge. From there, we’ll work with you to sharpen the scope around the questions you find most interesting. These don’t have known answers: the aim is to make real progress on problems worth solving, and we expect your ideas to shape where the work goes.

Teams will then meet for check-ins through the rest of Michaelmas and Hilary, with support along the way. Compute credits are included for all teams. The programme culminates in a final showcase at the end of Hilary, where teams will present what they’ve built, discovered, or proved, with a total prize pot of £5,000. 

Students at hackathon

The Challenges

  1. Spec-Driven Agentic Systems

A team builds a minimal spec-driven agent architecture in which a YAML/JSON specification describes a small ensemble of agents, their interfaces, and the tasks they handle. Changes to the spec trigger automated impact analysis identifying which agents and prompts are affected, regenerate those components, and run a self-test suite before redeployment. The deliverable is a working reference system of four to six agents plus a live demonstration of a spec change propagating end-to-end.

  1. Prompt Compression and Meta-Tuning for Agentic Workflows

A team assembles a benchmark of long agentic prompts (instructions, demonstrations, tool histories, memory) and implements two or three compression strategies — LLM-summarisation, learned token deletion, and structured retention of demonstration fragments. They evaluate task fidelity against prompt length on a standard agentic benchmark and build an interactive viewer showing which prompt regions compress safely and which break the task.

  1. Numeric Precision and Subgraph Simplification for ML Systems

A team picks one of two sub-problems. Either (a) train a small open model at FP8/BF16/FP16 across architectures, measure stability, and build a diagnostic that predicts instability from weight distributions; or (b) build a tool that proposes equivalent-but-faster subgraph rewrites for an open model family and verifies correctness against a held-out test set.

  1. LLM Robustness

A team picks one robustness axis — tokenisation sensitivity, generalisation dynamics under scale, or adversarial fine-tuning attacks — and runs a focused empirical study on one or two open models. They produce a small reusable benchmark suite and one clear empirical claim.

  1. Verifiable Trust Delegation in Multi-Agent Systems

A team implements a token-exchange middleware for one open agent framework (LangGraph, AutoGen, or CrewAI) that bounds delegated authority across multi-hop agent invocations. They define a small capability lattice, implement the protocol, verify one non-amplification property using ProVerif or a hand-proof, and benchmark latency overhead at one-, three-, and five-hop depths.

  1. Capability Containment and Cascade Failure Prevention

A team builds a multi-agent sandbox with configurable per-agent capabilities and implements a circuit-breaker that detects anomalous invocation patterns and isolates misbehaving agents. They run fault-injection scenarios — compromised agent, runaway invocation loop, data dependency cascade — and produce a live dashboard visualising blast radius before and after containment.

  1. Just-in-Time Access Control / Zero-Standing-Privilege

A team implements a JIT credential issuance system for a small service-oriented demo (three to five services), measures per-request latency overhead, formally analyses one security property using a lightweight tool such as ProVerif, and characterises the availability impact of credential-issuance unavailability.

  1. Observability of AI-Assisted Developer Productivity

A team builds a measurement framework that handles the full observability spectrum of AI coding tools, from visible autonomous-agent commits to invisible inline completions. They pilot the framework on a public mixed-attribution code corpus and produce an empirical write-up of what can and cannot be inferred about AI contribution from observable signals, with a small dashboard.

  1. Latency-Aware Scheduling for Heterogeneous LLM Inference

A team builds a small inference scheduler serving mixed workloads — interactive queries, batch summarisation, multiple model sizes, multiple SLA tiers — over one or two open LLMs. They implement a FIFO baseline plus a length-aware policy and measure head-of-line blocking and SLA hit rates under realistic load.

  1. Checkpoint-Aware Gang Scheduling

A team builds a simulator (no real GPUs required) for distributed training jobs under stochastic hardware failure, implements a baseline gang scheduler and a checkpoint-aware variant that uses checkpoint age in victim selection, runs trace-driven evaluation against published cluster traces (Alibaba, Google), and reports the goodput delta.

  1. Predictive Data Placement in Storage Hierarchies

A team builds a temporal point process model of file access patterns from a public HPC storage trace, integrates it with a tier-migration policy, and simulates predictive placement against a static rule-based baseline. They measure hot-data hit rate, migration overhead, and the dollar-equivalent cost saving at petabyte scale.

  1. Incremental Build Correctness Verification

A team builds a runtime verifier for incremental builds in a Bazel or Buck2 monorepo that detects under-approximated dependencies via build-action instrumentation. They evaluate on an open-source monorepo with planted spurious-edge pruning and access-restricted graph partitions, reporting how many silent correctness failures the verifier catches.

  1. Adaptive Query Isolation for HTAP Workloads

A team sets up CockroachDB or TiDB running an HTAP workload (CH-benCHmark, or TPC-C combined with TPC-H) and implements a query-plan-aware admission control mechanism that throttles analytical load when transactional latency exceeds a target. They measure transactional SLA preservation under increasing analytical pressure.

  1. Knowledge Graph Temporal Drift Detection

A team builds a drift detector over a public temporal knowledge graph (YAGO or Wikidata) plus a synthetic enterprise-style graph with planted drift. They implement temporal graph embeddings and a freshness score, and produce a viewer that highlights stale, contradictory, or orphaned subgraphs requiring attention.

  1. Gamified Simulation for Agentic Organisation Management

A team builds a simulation in which the player manages an organisation containing both human and AI agents, allocating tasks, designing escalation paths, setting quality thresholds, and responding to agent failures. Three to four progressively complex scenarios are evaluated against an expert-elicited optimal policy. Showcase audiences can play a five-minute scenario live.

  1. Detecting Misrouted Instructions in Hierarchical Agent Systems

A team builds a small hierarchical agent demo, injects controlled misrouting (capability mismatch, context mismatch, semantic misinterpretation), and implements two detectors — a capability-profile consistency check and an output anomaly detector using contrastive scoring against what a correctly-routed agent would produce. They measure precision, recall, and detection latency against the injected ground truth.

With thanks to G-Research

G-Research is a leading quantitative research and technology company, specialising in developing cutting-edge models and software for financial markets. Their commitment to advancing machine learning research and supporting future talent underscores their dedication to student entrepreneurship and innovation. 

We are grateful to G-Research for providing the challenges, mentorship and support that make the AI/ML Challenges possible.

Hackathon