Generated from runtime authorities
Capability coverage
Current Heliox source declares 157 action types and registers concrete providers for all 157 through 21 executable specialists. Availability still depends on operating system, permissions, credentials, hardware, and third-party services.
Important limitation
Verification depth
Provider coverage is not the same as independent outcome verification. Eighteen actions currently have a separate observed post-condition verifier; 139 rely on the executor result. Heliox must not describe all 157 actions as independently verified.
Live automation
Cross-platform CI
Python, frontend, visual regression, and Rust checks run across the committed CI matrix. Marketplace moderation and installer builds have separate gates.
Test certificate validated
Windows signing
The SignPath test-policy workflow has signed and verified the Windows EXE, MSI, and embedded application. The production certificate remains pending, so current public installers must not yet be described as production-signed.
Measured guarded path · 100 iterations
Latency distributions, not one headline number
The ready-daemon CPU-usage path measured 26.476 ms median, 27.999 ms p95, and 28.887 ms p99 on Windows/Python 3.12.6 with zero model calls. The path includes local planning, routing, risk assessment, real execution, post-condition verification, and response shaping.
59/59 regression cases
Intent routing controls
The deterministic dispatch corpus covers URLs, browser clicks, screen analysis, status queries, app launches, file reads, forensics, a browser workflow, and ambiguous phrases that must reach model planning. Median routing latency was 0.021 ms. This fixed corpus is not population-level language accuracy.
Shared-loop protection
Concurrent responsiveness
A real one-second CPU monitor allowed 66 concurrent scheduler heartbeats; the maximum gap was 16.575 ms on Windows. This protects voice, camera, approval, and WebSocket work from a one-second event-loop freeze. It is not UI or hardware-input latency.
Process-isolated local speech
Heavy voice memory is released
A real Kokoro file-synthesis run measured 21.401 seconds cold and 0.138 seconds warm. The parent retained zero Torch/CUDA modules, and the speech worker exited after its 10-second idle window. This does not measure audible quality, speaker compatibility, or universal latency.
Subscription-backed planning · planning only
3/3 fixed planning cases
One developer-machine Codex CLI run passed three fixed health, browser, and evidence-first planning prompts at 14.708 seconds median. No proposed action was executed and every case contained zero destructive actions. This does not establish Claude behavior, universal provider availability, or runtime outcomes.
Calibrated learned-risk evidence
World model: useful, bounded, inspectable
The shipped coarse risk predictor records 36,000 training and 5,400 temporal-validation samples across 12 action types. Held-out MAE improved 54.0154% for disk delta and 99.4124% for process delta versus a zero predictor; 5/5 direction invariants passed. Deterministic policy remains authoritative, and this is not a general physical-world or user-intent model.
Software path tested
Voice, gesture and gaze
Routing, cancellation, calibration, geometry, temporal filtering, and fusion have automated coverage. Accuracy across microphones, cameras, lighting, backgrounds, languages, accents, and users still requires human hardware testing.
Research boundary
Neural intent
Synthetic BrainFlow and recorded EEG playback paths are implemented. Heliox has not established live headset control accuracy, medical validity, or clinical use. It must be described as recorded/synthetic EEG research until live evidence exists.
Known limits remain visible
What this evidence does not prove
- Universal compatibility with websites, applications, operating systems, peripherals, or third-party APIs.
- Physical camera, microphone, speaker, accessibility-permission, EEG, or human-factors accuracy.
- That snapshots can reverse messages, purchases, remote operations, pushed commits, or every external effect.
- That learned-risk or world-model output can override deterministic safety policy.
- That the currently published Windows installers carry a production-trusted certificate.