Runs ~20 options strategies against live market data in shadow mode, records every hypothetical fill under worst/base/optimistic assumptions, and grades each with anytime-valid e-processes. Places no orders.
-
Updated
Jul 22, 2026 - Python
Runs ~20 options strategies against live market data in shadow mode, records every hypothetical fill under worst/base/optimistic assumptions, and grades each with anytime-valid e-processes. Places no orders.
Repository for the paper "Auditing Pay-Per-Token in Large Language Models", AISTATS'26
This repository contains the code for the paper "Optimizing Social Utility in Sequential Experiments".
A kernel-userland protocol enforcing information-theoretic bounds on AI adaptivity leakage, benchmark gaming, and capability spillover.
Anytime-valid sequential auditing of disclosed loot-box drop rates
Research implementation of anytime-valid group-invariance tests with kernel, adaptive Gaussian, and neural statistics.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
Label-free drift attribution (model / noise / world / annotator) with anytime-valid guarantees and a frozen pre-registration — closes the zero-label misattribution gap 0.50→0.00 on the recoverable cause; negatives reported as measured boundaries.
Publicly verifiable proof that an LLM endpoint actually spent the compute you paid for — zero provider cooperation. Closes the effort gap (Hollow-LLM) and the transferability gap (IRIS).
Benchmark for statistically valid AI scientist systems, using audit-closed protocols, transparency logs, and sequential inference to prevent false discoveries in autonomous research agents.
Add a description, image, and links to the e-values topic page so that developers can more easily learn about it.
To associate your repository with the e-values topic, visit your repo's landing page and select "manage topics."