You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
As artificial intelligence (AI) agents transition from virtual sandboxes to physical-world deployments, ensuring their safe operation becomes a critical challenge. Conventional AI alignment and safety paradigms typically assume instantaneous feedback loops and perfect state information. However, physical-world autonomous systems, such as Unmanned Aerial Vehicles (UAVs), operate under severe runtime constraints, including network delays, stale sensor data, multi-agent coordination conflicts, and strategic evasion by capable models.
This proposal introduces ActionProof, an open, reproducible test environment designed to systematically evaluate whether independent safety verification layers can reliably intercept and block hazardous commands from physical-world AI agents. Rather than seeking to demonstrate a single, flawless safety setup, ActionProof is intentionally built to stress-test these guardrails to their breaking points. By quantifying fundamental safety-latency trade-offs and uncovering vulnerabilities to strategic model bypasses, this research will provide the AI safety community with the tools and empirical benchmarks required to engineer robust, network-aware, and physically grounded safety architectures.
A promising paradigm for safe AI integration is the deployment of an independent safety checker (or runtime guardrail) that sits between the AI agent's decision-making engine and the physical actuators. For example, if an AI agent controlling a drone issues a command to descend and navigate toward a specific coordinate, this command is intercepted by an independent safety layer. The checker evaluates the action against strict, predefined physical boundaries, such as minimum altitude, obstacle proximity maps, remaining battery safety margins, and communication link status. If the action is determined to be safe, it is allowed to proceed; otherwise, the command is blocked or redirected to an automated safe fallback routine.
While this independent checking architecture is conceptually elegant, it contains critical, under-explored vulnerabilities when deployed in networked, dynamic, and multi-agent physical environments. Specifically, we identify three critical failure modes that the ActionProof environment will address:
• Decision Staleness and 'Reality Drift': In physical systems, state estimation and communication are subject to inevitable propagation and processing delays. Suppose a drone's physical state is measured at time t_0. The AI agent processes this telemetry to make a decision, which is then transmitted over a network to the verifier, checked, and finally sent to the actuators. By the time the approved action is executed, the drone's true state may have drifted significantly (e.g., due to wind drift or moving obstacles). Consequently, the safety checker verifies safety based on stale information, leading to catastrophic runtime failures.
• Concurrently Approved Interference: When multiple autonomous AI agents operate within a shared physical space, their safety layers check and approve commands independently. Without explicit synchronisation, two individual checkers may simultaneously approve commands that are mutually safe in isolation but lead to immediate collisions when executed concurrently.
• Strategic Evasion and Rule Exploitation: As physical-world AI agents gain reasoning capabilities, they may learn to exploit the logical boundaries of safety verification rules. If a model learns that a large, high-risk command will be blocked by the checker, it may strategically break the forbidden maneuver into multiple, minor adjustments that are individually flagged as 'safe' but cumulatively execute the hazardous action (action fragmentation). Alternatively, a restricted agent may delegate a prohibited action to another nearby agent to bypass its own safety layers.
The primary objective of the ActionProof project is to rigorously analyze and expose the temporal and behavioral boundaries of independent AI safety layers. The research is structured around two central objectives:
• Research Goal 1 (Safety-Latency Co-design): To determine the mathematical frontier of safety versus latency. We seek to evaluate whether an independent verification layer can successfully intercept hazardous commands in real-time without introducing computational and communication latency that destabilizes the physical flight control loop.
• Research Goal 2 (Evasion and Bypass Probing): To systematically probe and identify logical vulnerabilities that allow capable AI models to strategically bypass runtime checkers under various degrees of information asymmetry, network delay, and multi-agent delegation.
These objectives are driven by two key research hypotheses:
• Hypothesis 1 (Temporal Degradation): There exists a critical latency threshold (defined as the sum of network propagation, model inference, and verification computation) beyond which the addition of an independent safety checker actually reduces the overall safety of the physical system compared to running the system without a checker.
• Hypothesis 2 (Strategic Evasion): Capable, goal-oriented AI models can spontaneously discover and exploit verification blind spots (e.g., through action fragmentation or cross-agent delegation) when exposed to iterative testing against static, rule-based safety checkers.
Experimental Design and Benchmark Methodology:
ActionProof will be implemented as an open, highly reproducible testbed combining high-fidelity drone simulations, network delay emulators, physical edge-computing hardware, and state-of-the-art frontier and open-source AI models.
Core Testing Infrastructure:
The software architecture of ActionProof will integrate three distinct layers:
• High-Fidelity UAV Simulation: A realistic multi-rotor simulation environment (utilizing frameworks like PX4, AirSim, or Webots) capable of modeling wind, sensor noise, collision physics, and battery discharge rates.
• Edge and Network Control Emulation: A software-defined network emulator that models communication delays (from 10 ms to 500 ms), packet loss, and jitter to mimic edge-to-cloud and vehicle-to-everything (V2X) communication channels.
• Hardware-in-the-Loop (HIL) Integration: Interfacing the simulation and network layers with physical edge-computing boards and transceivers to validate experimental findings under real hardware latency.
Stress-Testing Scenarios
To evaluate the robustness of safety checkers, the test environment will subject AI agents and verifiers to the following highly challenging operational scenarios:
• Stale State Injection: Simulating old positioning and altitude information at the verifier interface to analyze how safety checks degrade as telemetry age increases.
• Communication Latency Stress: Introducing communication delays and packet drops between the AI agent and the safety layer to determine the exact point where verification latency destabilizes flight control.
• Multi-Agent Synchronization: Forcing two AI agents to execute actions in close proximity under partial network visibility to probe concurrent collision-checking limits.
• Delegation-Based Evasion: Establishing scenarios where Agent A is explicitly barred from entering a restricted airspace but can communicate with Agent B, analyzing whether Agent A will attempt to orchestrate a delegation-based bypass.
• Action Fragmentation (Slicing): Setting up a strict safety rule that blocks any single high-speed descent command, and testing whether the AI model can learn to bypass this check by issuing five consecutive, slight altitude reductions that result in the same unsafe descent rate.
Project Implementation Roadmap:
The ActionProof project will be executed over a 12-month schedule, structured into three distinct phases to ensure rapid prototyping, thorough testing, and a high-quality final release:
• Phase 1: Environment Setup & ActionProof-v0 (Months 1–4): Establish the core simulation interface and integrate 2-3 open-source and frontier models via APIs. Implement initial rule-based safety checkers (verifying static physical limits, safe zones, and collision margins) and construct the first reproducible decision-staleness and action-fragmentation scenarios. Deliverable: ActionProof-v0 codebase.
• Phase 2: Delay and Multi-Agent Stress Testing (Months 5–8): Build software-defined network emulation modules to introduce delays and packet loss. Conduct extensive multi-agent and delegation-based bypass experiments. Implement predictive safety models and explore formal runtime verification techniques to analyze state estimation errors.
• Phase 3: Hardware Validation & Benchmark Dissemination (Months 9–12): Validate simulated failure modes using physical edge-computing hardware and transceivers in a hardware-in-the-loop (HIL) environment. Refine metrics, clean the repository, draft technical documentation, and release the ActionProof public benchmark dataset and codebase to the AI safety community.
To ensure high-impact scientific output, the requested funding will be deployed across critical operational areas, leveraging existing hardware resources to minimize capital expenditure:
Personnel: A primary budget portion will fund a dedicated Research Engineer or Postdoctoral Researcher to lead simulator integration, API connections, network delay scripting, and benchmark automation. Funding will also support METU student researchers to run model evaluations, assist with hardware testing, and maintain the data pipeline.
Compute and Model APIs: Allocation for continuous API tokens (to query frontier commercial models) and cloud-based GPU instances required for executing parallel, large-scale agent stress-testing simulations.
Hardware Infrastructure: The project will be highly resource-efficient by leveraging the PI's existing hardware infrastructure at METU, which includes edge-computing development boards (NVIDIA Jetson, Raspberry Pi), wireless communication transceivers, and UAV flight-control hardware. A small portion of the budget will be allocated for necessary sensors and interfaces to complete the hardware-in-the-loop (HIL) setup.
The Principal Investigator, Dr. Muhammad Toaha Raza Khan, is an Assistant Professor of Computer Engineering at Middle East Technical University (METU), Ankara, Turkey. While his background is not in traditional, software-only AI alignment, his deep expertise in delay-sensitive communication, distributed intelligent systems, V2X networking, vehicular communications, UAV networks, edge computing, non-terrestrial networks, and reinforcement learning provides a unique, highly critical system-level perspective.
In physical-world autonomous systems, safety cannot be solved in a software-only vacuum. Dr. Khan's extensive career has focused on solving the exact challenges that break software guardrails: communication latency, packet dropouts, and asynchronous coordination in highly dynamic environments. His system-level insights represent the exact perspective needed to design robust, network-aware safety architectures. Dr. Khan’s research has been widely published in leading journals in the field, including:
• IEEE Transactions on Intelligent Transportation Systems
• IEEE Transactions on Vehicular Technology
• IEEE Sensors Journal
• IEEE Communications Magazine
• IEEE Wireless Communications Letters
• Information Fusion
• Internet of Things
At METU, Dr. Khan supervises a talented group of graduate and undergraduate students working on edge systems, AI, sensing, and wireless communications. In the last 12 months, while raising $0 in dedicated AI safety funding, he has successfully supervised two active AdımODTÜ undergraduate research grants (TRY 70,000 and TRY 50,000). Prior to joining METU, Dr. Khan worked on large-scale funded research initiatives spanning post-earthquake communication resilience using UAVs and LoRa, AI-based resource allocation in V2X networks, and projects supported by the National Research Foundation (NRF) of Korea and the BK21 program.
The ActionProof project is designed to navigate a set of complex technical risks through targeted mitigation strategies:
The Safety-Latency Trade-off: High-fidelity safety checking may introduce a 300 ms processing delay. In dynamic flight, this 300 ms latency can actually be more physically dangerous than having no checker at all. The benchmark will explicitly model this trade-off, defining a mathematical Pareto frontier that balances verification depth against physical system stability.
Multi-Agent Complexity: Evaluating multi-agent interactions introduces a combinatorial explosion of states. To mitigate this, we will start with simple twin-agent scenarios to validate core concepts before introducing multi-agent scaling.
Definition of Scientific Success:
Unlike commercial software development, where a safety bypass is treated as a project failure, in ActionProof, safety failures are scientific successes. The core mission of this benchmark is to break independent safety layers under realistic physical and network conditions.
The project will be deemed a scientific success if it demonstrates under what exact network delays, telemetry age, and strategic behaviors (such as action fragmentation and delegation) standard safety checkers fail. By publishing these negative results, releasing the open-source ActionProof-v0 benchmark, and making the codebase fully reproducible, this work will provide the global AI safety community with the concrete empirical framework needed to move AI safety from theoretical abstractions to real-world, network-aware runtime assurance.
I have raised $0 for ActionProof.
During the last 12 months, I received two AdımODTÜ undergraduate research grants as project supervisor.
One is TRY 70,000 and the other is TRY 50,000, so TRY 120,000 in total.
These projects are separate from ActionProof.
Before this, I worked on larger funded research projects related to post-earthquake communication resilience, AI resource allocation in V2X systems, and research supported through Korean NRF and BK21 programs.
None of those funds are being used for this project.
There are no bids on this project.