- Spec-Driven Agentic Systems
A team builds a minimal spec-driven agent architecture in which a YAML/JSON specification describes a small ensemble of agents, their interfaces, and the tasks they handle. Changes to the spec trigger automated impact analysis identifying which agents and prompts are affected, regenerate those components, and run a self-test suite before redeployment. The deliverable is a working reference system of four to six agents plus a live demonstration of a spec change propagating end-to-end.
- Prompt Compression and Meta-Tuning for Agentic Workflows
A team assembles a benchmark of long agentic prompts (instructions, demonstrations, tool histories, memory) and implements two or three compression strategies — LLM-summarisation, learned token deletion, and structured retention of demonstration fragments. They evaluate task fidelity against prompt length on a standard agentic benchmark and build an interactive viewer showing which prompt regions compress safely and which break the task.
- Numeric Precision and Subgraph Simplification for ML Systems
A team picks one of two sub-problems. Either (a) train a small open model at FP8/BF16/FP16 across architectures, measure stability, and build a diagnostic that predicts instability from weight distributions; or (b) build a tool that proposes equivalent-but-faster subgraph rewrites for an open model family and verifies correctness against a held-out test set.
- LLM Robustness
A team picks one robustness axis — tokenisation sensitivity, generalisation dynamics under scale, or adversarial fine-tuning attacks — and runs a focused empirical study on one or two open models. They produce a small reusable benchmark suite and one clear empirical claim.
- Verifiable Trust Delegation in Multi-Agent Systems
A team implements a token-exchange middleware for one open agent framework (LangGraph, AutoGen, or CrewAI) that bounds delegated authority across multi-hop agent invocations. They define a small capability lattice, implement the protocol, verify one non-amplification property using ProVerif or a hand-proof, and benchmark latency overhead at one-, three-, and five-hop depths.
- Capability Containment and Cascade Failure Prevention
A team builds a multi-agent sandbox with configurable per-agent capabilities and implements a circuit-breaker that detects anomalous invocation patterns and isolates misbehaving agents. They run fault-injection scenarios — compromised agent, runaway invocation loop, data dependency cascade — and produce a live dashboard visualising blast radius before and after containment.
- Just-in-Time Access Control / Zero-Standing-Privilege
A team implements a JIT credential issuance system for a small service-oriented demo (three to five services), measures per-request latency overhead, formally analyses one security property using a lightweight tool such as ProVerif, and characterises the availability impact of credential-issuance unavailability.
- Observability of AI-Assisted Developer Productivity
A team builds a measurement framework that handles the full observability spectrum of AI coding tools, from visible autonomous-agent commits to invisible inline completions. They pilot the framework on a public mixed-attribution code corpus and produce an empirical write-up of what can and cannot be inferred about AI contribution from observable signals, with a small dashboard.
- Latency-Aware Scheduling for Heterogeneous LLM Inference
A team builds a small inference scheduler serving mixed workloads — interactive queries, batch summarisation, multiple model sizes, multiple SLA tiers — over one or two open LLMs. They implement a FIFO baseline plus a length-aware policy and measure head-of-line blocking and SLA hit rates under realistic load.
- Checkpoint-Aware Gang Scheduling
A team builds a simulator (no real GPUs required) for distributed training jobs under stochastic hardware failure, implements a baseline gang scheduler and a checkpoint-aware variant that uses checkpoint age in victim selection, runs trace-driven evaluation against published cluster traces (Alibaba, Google), and reports the goodput delta.
- Predictive Data Placement in Storage Hierarchies
A team builds a temporal point process model of file access patterns from a public HPC storage trace, integrates it with a tier-migration policy, and simulates predictive placement against a static rule-based baseline. They measure hot-data hit rate, migration overhead, and the dollar-equivalent cost saving at petabyte scale.
- Incremental Build Correctness Verification
A team builds a runtime verifier for incremental builds in a Bazel or Buck2 monorepo that detects under-approximated dependencies via build-action instrumentation. They evaluate on an open-source monorepo with planted spurious-edge pruning and access-restricted graph partitions, reporting how many silent correctness failures the verifier catches.
- Adaptive Query Isolation for HTAP Workloads
A team sets up CockroachDB or TiDB running an HTAP workload (CH-benCHmark, or TPC-C combined with TPC-H) and implements a query-plan-aware admission control mechanism that throttles analytical load when transactional latency exceeds a target. They measure transactional SLA preservation under increasing analytical pressure.
- Knowledge Graph Temporal Drift Detection
A team builds a drift detector over a public temporal knowledge graph (YAGO or Wikidata) plus a synthetic enterprise-style graph with planted drift. They implement temporal graph embeddings and a freshness score, and produce a viewer that highlights stale, contradictory, or orphaned subgraphs requiring attention.
- Gamified Simulation for Agentic Organisation Management
A team builds a simulation in which the player manages an organisation containing both human and AI agents, allocating tasks, designing escalation paths, setting quality thresholds, and responding to agent failures. Three to four progressively complex scenarios are evaluated against an expert-elicited optimal policy. Showcase audiences can play a five-minute scenario live.
- Detecting Misrouted Instructions in Hierarchical Agent Systems
A team builds a small hierarchical agent demo, injects controlled misrouting (capability mismatch, context mismatch, semantic misinterpretation), and implements two detectors — a capability-profile consistency check and an output anomaly detector using contrastive scoring against what a correctly-routed agent would produce. They measure precision, recall, and detection latency against the injected ground truth.