Takahashi, K
An interface theory that ensures genuine stable improvement in self-improving systems even when the evaluation criteria change endogenously.
When AI systems improve their own code or policies, the evaluation criteria (benchmarks, parsers, routing logic, etc.) also change, making it difficult to distinguish short-term gains from genuine improvement. Existing research mainly addresses self-improvement under fixed evaluation criteria, but the stability problem when the criteria themselves change has not been established.
The author separates the self-improving system into four layers (base admissibility, certified stable gain, debt-bounded stability, governance safety) and mathematically defines the stability of each layer through replayable observable interfaces. Specifically, typed-noise comparator certificates, replayable state lifts, lower-oracle stable-gain intervals, delayed shadow certification, contradiction obligations, semantic-retention margins, proof-carrying escalation lanes, and unresolved-downside debt reserves are introduced to derive stability conditions.
The main result shows that stability and growth claims must be stated conditionally on published replayable instrumentation and conservative envelope objects. It also derives small-gain style stability wrappers for the coupled improver–evaluator–memory–queue loop, certified window contraction conditions, replayable lower certificates under delayed audit, and comparison theorems for explicit regime–policy pairs. In particular, it shows that governance edits can be frontier-expanding investments even when they do not provide immediate production capability increments, provided the replay, challenge, contradiction preservation, escalation, proof, and verification interfaces are sufficiently robust.