SKILL-ROLE: Extreme Testing Rigor for Applications & Games
1. Core Philosophy & Mandates
This skill configures the developer or agent to architect, implement, and verify general software systems—including interactive applications, engines, and game codebases—with extreme testing rigor inspired by aeronautical-grade and safety-critical standards [4]. The system design must be governed by three non-negotiable mandates:
- "If it isn't tested, it doesn't work" [4]. No logic branch, feature, or correction is complete without comprehensive, multi-layer verification.
- "Testability must be an architectural objective" from day one [24]. Code must be structured with strict, decoupled boundary interfaces that permit deep mocking, automated fault injection, and simulated platform abstractions [7, 8, 10].
- "Fly what you test, test what you fly" [7]. The code delivered to production must remain structurally identical to the code that underwent validation [7]. Testing hooks, defensive assertions, and simulated failure wrappers must either remain safely embedded in the codebase or be strictly controlled under identical runtime compilation profiles [10].
2. Multi-Plane Verification Architecture
Testing must be organized across three distinct verification planes to catch failures at different levels of the application runtime lifecycle [6]:
+-------------------------------------------------------------------------+
| THREE VERIFICATION PLANES |
+-------------------------------------------------------------------------+
| 1. LOGICAL STATE & COVERAGE | 2. RUNTIME SANITIZERS | 3. TARGET PLATFORM |
| | | |
| - Requirement analysis | - Memory leaks / safety| - Same compiled |
| - 100% MC/DC branch check | - Async / race audits | production code |
| - Structural invariants | - Leak tracking | - Physical devices|
| - State coverage matrix | - Out-of-bounds guards | and target OS |
+-------------------------------------------------------------------------+
Plane 1: Verify Logical State & Coverage
- 100% Branch Coverage: Ensure every logical conditional statement is fully exercised across all possible execution paths [4, 5].
- 100% MC/DC (Modified Condition/Decision Coverage): Every complex decision containing multiple sub-conditions must be executed such that each sub-condition is proven to independently affect the decision's outcome [4, 5].
- Deep Boundary Integrity: Instrument automated tests to target the exact boundaries of numeric, physics, array, or transition limits (e.g. at $X-1$, $X$, and $X+1$) [6].
- Test-to-Code Volume: The verification suite must follow a strict volume ratio where the test harness and test case logic are significantly larger than the application source code itself [24]. Roughly 10% to 20% of production source code should be composed of internal diagnostics, simulated interfaces, and diagnostic hooks [20, 24].
Plane 2: Verify Runtime Safety & Sanitizers
- Run the code under strict compiler/interpreter sanitizers or runtime
environment trackers:
- Memory & Resource Leak Trackers: Ensure no heap allocations, timers, threads, textures, or listeners are leaked [6].
- Async & Thread Sanitizers: Check for data races, state race conditions, and deadlocks in multi-threaded/concurrent code [6].
- Type & Bounds Checkers: Trap array out-of-bounds, uninitialized variables, integer overflows, and null-pointer dereferences [6].
Plane 3: Verify on Target Parity
- Execute tests on the exact target hardware platforms (e.g., target mobile devices, gaming consoles, specific OS architectures) using production-optimized build profiles to catch compiler optimizations or platform-specific bugs [6, 18].
3. Inline Instrumentation & Defensive Checking
Code must be actively instrumented with defensive checking patterns to document, check, and prove state assumptions during execution [13, 16].
A. State Invariants (assert(X))
- Every critical class, service, or state machine must define a function to check its structural state invariants [13].
- These assertions serve as executable comments describing the absolute truth of the program state [13].
- Example Invariant Validation: A game's spatial grid component should assert that no entity is registered in a coordinate cell without having its internal coordinates matching that exact cell [14].
- Invariants must be active during testing, debugging, and coverage runs, but compiled to low-overhead or no-op paths in final public builds if maximum performance recovery is required [14, 16].
B. Defensive Branch Tracking (ALWAYS(X) & NEVER(X))
- Use conditional wrappers to document and check paths that are considered
mathematically or logically "impossible" under normal operation but are kept
for defensive safety [15, 16]:
ALWAYS(X)evaluatesXand triggers a runtime panic or test failure if it ever evaluates to false [15, 16].NEVER(X)evaluatesXand triggers a panic or test failure if it ever evaluates to true [15, 16].
- During testing coverage analysis, these helper structures return constant
values (e.g.,
ALWAYS(X) => 1,NEVER(X) => 0) to allow coverage reporting engines to bypass defensive branches that cannot be naturally triggered by standard test cases [15, 16].
C. Boundary Trapping (testcase(X))
- Inject
testcase()logging markers directly alongside condition boundaries so testing harnesses can programmatically verify that extreme boundary parameters have been hit and evaluated [6, 16].
4. Component Mocking & Systematic Fault Injection
A robust application must not assume that downstream subsystems, network layers, hardware resources, or input peripherals behave perfectly. The system must decouple components via dependency injection interfaces to systematically simulate catastrophic environments [7, 8].
+--------------------+
| Testing Engine |
+--------------------+
|
v (Controls injected faults) [10]
+--------------------+ +------------------------------+
| Application/Game | <-----> | Simulation & Fault Control | (Diagnostic Control System) [10]
| Core Logic | +------------------------------+
+--------------------+
|
v (Calls decoupled interface)
+--------------------+
| Decoupled Mock | <--- Simulate: Resource loading fails [7], packet loss [7],
| Boundary Interface| state corruption [10], out-of-memory errors [8]
+--------------------+
A. Systematic Resource & Memory Allocation Failures
For critical operations (such as loading a gameplay level, executing an asset transfer, or processing a transaction), mock the underlying system allocator or resource loading system to test recovery logic [8]:
- The Single-Failure Loop:
- Set the mocked component or allocator to fail exactly once on the $N$-th call, for $N = 1, 2, 3, \dots$ [8].
- Run the full test execution.
- Increment $N$ and repeat the test run until the operations complete fully with zero simulated allocation failures [8].
- The Persistent-Failure Loop:
- Set the mocked component to fail all subsequent operations starting at the $N$-th call, for $N = 1, 2, 3, \dots$ [8].
- Run the test execution.
- Increment $N$ and repeat the sequence until the system runs to completion [8].
- Verification Criterion: Across every single simulated failure, the application under test must gracefully clean up all resources, leak zero memory, degrade gracefully, and return proper, predictable error states without crashing.
B. Input, Network, & State Interruption Mocking
- Simulate Network/Packet Faults: Mock the network sockets or service layer to drop, delay, or duplicate packets on the $N$-th API call [9]. Verify that state replication models do not fall out of sync.
- State Snapshot Corruption (Save State Recovery):
- Configure the mock state-serializer to take a snapshot of the runtime state on the $N$-th update [9].
- Corrupt the Snapshot: Apply random mutations, partial writes, or garbage values to the snapshot to emulate a crash, write tearing, or incomplete save-data serialization [10].
- Verification: Restart the application core directly on the corrupted state snapshot [10]. Verify that recovery systems successfully repair, safely ignore, or gracefully reject the corrupt data, resetting the state to a safe default without crashing or creating inconsistent states [1, 10].
- Repeat this mutation sequence multiple times across different execution points $N$ [10].
5. Continuous Automated Fuzzing & Logical Inconsistencies
A. Coverage-Driven Action Fuzzing
- Execute automated fuzzing engines or randomized action generators (e.g., automated headless agents executing randomized player actions, UI clicks, or API payloads) [21].
- Maintain record of active code coverage pathways during fuzzing [21].
- Retain only the input permutations or action sequences that activate brand-new logic branches to continuously expand the test suite's structural coverage [21].
B. Metamorphic / Relational Inconsistency Detection
- Apply relational consistency rules to detect logical bugs in queries, search engines, UI render trees, or game state updates [22, 23].
- The Logical Equivalence Rule:
- Let $A$ be a complex action, search, or update on a data state that produces output $Y$ [22].
- If we perform a structurally filtered or segmented variation of action $A$ (e.g. running the search, and then applying a manual client-side filter post-search), the filtered output must yield the exact same result $Y$ [22, 23].
- Automate tests that run these logically equivalent pipelines in parallel using random data generation. If their outputs differ, flag a logical bug [23].
6. Verification of the Verification (Mutation Testing)
To guarantee that the testing suite is actually capable of detecting logical regressions, developers must run periodic mutation testing on the codebase [17].
- Execution: Introduce small, single-line modifications to the code logic
under test [17]:
- Invert a boolean conditional check (e.g. from
<to>=or!=to==) [17, 18]. - Convert a math operation or dynamic calculation into a constant value or no-op [17].
- Comment out a state transition or defensive handler [17].
- Invert a boolean conditional check (e.g. from
- Validation: Execute the testing suite.
- If the test suite fails, the test suite successfully detected the bug.
- If the test suite passes, the test suite has a gap [17]. A new test case must be added to cover that logical decision pathway.
- Shortcut Verification: Pay special attention to optimization gates (e.g.,
if (shortcutAvailable) { fastShortcut() } else { slowFallback() }) [18]. If mutatingshortcutAvailableto always return false does not cause tests to fail becauseslowFallback()is perfectly correct, developers must write explicit tests targeting the shortcut branch to ensure performance optimization regressions do not go unnoticed [18].
7. Operational Testing Checklist
Implementers and verification workflows must check off these items before declaring any code change ready:
- [ ] Structural State Invariants Placed: Define state invariants for every
dynamic component or state machine. Check them via
assert()[13, 14]. - [ ] Defensive Pathways Handled: Guard impossible code assumptions with
ALWAYS()/NEVER()checks [15, 16]. - [ ] Testcase Loggers Hooked: Mark boundary thresholds with
testcase()helper wrappers to verify boundary tests execute correctly [6, 16]. - [ ] Decoupled Boundary Interface Built: Decouple external APIs, inputs, memory allocation, and networking protocols into interfaces [7, 8].
- [ ] Systematic Single-Failure Loop Executed: Verify the system handles a failure at the $N$-th call gracefully without crashing or leaking memory [8, 9].
- [ ] Systematic Persistent-Failure Loop Executed: Verify the system behaves predictably when resource failures persist indefinitely [8].
- [ ] Corrupted State Recovery Verified: Snapshot the application state, inject mutations/corruptions, and verify recovery logic resolves gracefully [10].
- [ ] Test Suite Coverage Achieved: Run coverage utilities and ensure high-level branch/condition coverage [4, 6].
- [ ] Mutation Verification Checked: selectively patch code statements to verify the test suite fails when bugs are active [17].