Continuous companion loop
Listening, barge-in, actions, suggestions, risk interruption, and spoken follow-up remain coordinated.
Scroll to enter a living operating system where perception, agents, and safety move as one.
Control your entire computer with natural language, voice, and hand gestures. An open-source, privacy-first AI agent that plans, executes, and verifies complex multi-step tasks — running locally on your hardware.
Control your entire digital ecosystem with zero friction through four core modalities.
"Hey Heliox" ambient wake-word with VAD-based endpoints, barge-in support, and real-time push-free dispatch for frictionless execution.
Static poses (Palm, Pinch) and motion gestures (Two-Finger Swipe) with kinematic smoothing, 3D world-model, and optional gaze tracking fusion.
Continuous computer-vision loop taking screenshots every 3s, detecting active apps, and maintaining a rolling context buffer fed into the LLM planner.
Fire-and-forget background tasks that decompose, execute, and verify completely independent of the UI or main event loop.
Heliox OS combines reactive commands with opt-in proactive and background capabilities — all running through the same permission and verification pipeline.
Pattern-matches local screen context and surfaces visible, optional help before you ask. Accept/dismiss feedback tunes each pattern's timing and priority, and temporarily suppresses repeatedly rejected suggestions.
Independently reviews proposed plans, can warn, revise, or stop work that drifts from the request, accepts typed or spoken corrections while a task is running, and offers grounded next ideas after verified results.
Spawn complex multi-step background tasks that decompose, execute, and verify completely independent of the UI or main event loop.
Auto-bootstrapped computer vision tracks your contextual state cross-platform, natively bridging exactly what you see into the LLM planner.
"Hey Heliox" ambient wake-word with VAD-based speech endpoints and barge-in — start talking and Heliox stops mid-sentence to listen.
Static poses (Palm, Pinch) and motion gestures (Two-Finger Swipe) with kinematic smoothing. Opt-in 3D world-model via MediaPipe HandLandmarker and coarse gaze-tracking fusion.
Point to move the real OS cursor, pinch to click — opt in via Settings. Open palm always exits instantly. Off by default for safety.
On-device continual-learning loop personalizes pinch/thumb thresholds and wake-word matching from implicit signals — no retraining, bounded, resettable.
Animated, immersive Tauri overlays responding contextually to system actions with real-time visual feedback and status indicators.
A lightweight, dependency-free engine estimating attention, stress, and cognitive load from local interaction-history heuristics. Zero external models, no network, no license restrictions.
Real-time attention/stress/load estimates mapped onto the Svelte UI.
JARVIS automatically slows voice generation during high cognitive-load tasks.
High-risk actions evaluate cognitive stress first. If distracted, a 10-second auditory confirmation gate holds.
Subconscious background loop learns cognitive engagement patterns and encodes them to persona.md.
Notification pipeline buffers trivial alerts until a low-load resting state is detected.
Tasks predict aggregate cognitive demand. If executing risks exceeding mental bandwidth, JARVIS pauses.
Native intent fusion classifying spoken commands against current workload intensity for context-aware execution routing.
A true agentic system running a continuous ReAct loop with a modular multi-agent orchestrator — not a simple command runner.
LLM evaluates persistent memory context before reasoning
Converts natural language into structured multi-step action plan
Routes each action to 21 specialists across 20 source-scoped roles
21 domain specialists execute through reviewed OS, browser, and integration adapters
Post-execution verification confirms action success
Self-improvement engine learns from successes and failures
Five-tier permission system with confirmation gates & rollback
Complex goals auto-broken into dependency-aware subtask trees with parallel execution paths.
Pre-execution impact reports for dangerous commands showing risk level, affected files, and reversibility.
Successful reasoning chains stored and reused with keyword-indexed templates and success/failure rates.
Detached worktrees evaluated in a network-disabled, credential-free Docker runner — no automatic code promotion.
The planner schema declares 156 action types with 156/156 provider coverage. Availability depends on OS, dependencies, credentials, and security policy.
Two extension paths: auto-discovered local plugins with Ed25519 signature verification, and a reviewed GitHub marketplace with CI-enforced moderation.
Get real-time weather data for any location via voice or text commands.
Control Spotify playback, search tracks, and manage playlists via gestures or voice.
Bridge your smart home devices into the Heliox agent for unified automation.
Every plugin is signed. Unsigned, untrusted, or tampered packages are rejected before their manifest or code is loaded. Planner-triggered plugin actions route through the normal permission system.
Run entirely on local hardware for total privacy, or seamlessly connect to frontier cloud models. All AI outputs pass through structured schema validation before execution.
Read-only through root-level with confirmation gates and per-action granular approve/deny.
Pre-execution impact reports for dangerous commands with risk assessment and reversibility tracking.
Secondary LLM independently reviews Tier 3-4 plans before the confirmation gate fires.
Bundled MLP predicts action impact from training samples. Can interrupt but never remove rule-based warnings.
Btrfs/Timeshift on Linux, Windows Restore Points. If backend unavailable, destructive action doesn't run.
Free, CPU-only local speech model — no API key, no cost. Falls back to platform TTS if not installed.
Stop button genuinely kills in-flight subprocess (proc.kill()) instead of only stopping the next action.
API keys stored via platform keyring (GNOME Keyring / Windows Credential Manager). Never logged or sent to local LLMs.
The largest Heliox release yet: coordinated always-on voice, bounded adaptive learning, 21 specialists, 156 action types, opt-in gaze and 3D gesture paths, guarded neural-research workflows, and more truthful desktop execution.
Listening, barge-in, actions, suggestions, risk interruption, and spoken follow-up remain coordinated.
Verified experience, temporal memory, strategy evolution, and optional JEPA-style prediction advise plans without widening authority.
Concrete provider coverage spans desktop, browser, development, integrations, and bounded research workflows.
Voice, gaze, 3D hand tracking, cursor mode, and gesture workflows can run together with temporal false-positive rejection.
Synthetic BrainFlow and recorded EEGBCI evidence is available without misrepresenting it as proven live brain control.
Application resolution, browser targeting, approvals, cancellation, result reporting, and first-run setup are more deterministic.
Current main software validation
1,653 Python tests passed · 179 frontend tests passed · cross-platform visual and Rust checks passed · 0 high-severity npm vulnerabilities
Physical microphone, camera, gaze, gesture, audible TTS, and live-EEG accuracy remains device-dependent.
Measured software evidence · 13 August 2026
The public evidence bundle records distributions, regression controls, raw data, reproduction commands, and known exclusions. These are software-path measurements on one Windows host—not universal device, provider, network, or human-accuracy claims.
Guarded local path
28.640 ms
Median across 100 ready-daemon CPU-status requests; 30.490 ms p95 and zero model calls.
Intent controls
59 / 59
A fixed deterministic routing regression set, not population-level language accuracy.
Shared-loop health
65 ticks
Concurrent scheduler heartbeats during a real one-second CPU sample; 16.296 ms maximum gap.
Learned risk
36k / 5.4k
Training and temporal-validation samples for a bounded disk/process predictor; deterministic policy remains authoritative.
What these numbers exclude
Model-provider and network latency, browser page loads, UI rendering, microphone capture, audible TTS, camera, gesture, gaze, live EEG, and human accuracy all require separate evidence.
Download the compiled desktop app or build from source for development.
Download Heliox-OS_0.11.1_x64-setup.exe or the .msi from the latest published release.
Install and open Heliox OS.
Enter your API key (Gemini, OpenAI, Claude, Meta) in the Settings tab.
Download for Apple Silicon: Heliox-OS_0.11.1_aarch64.dmg
Download for Intel: Heliox-OS_0.11.1_x64.dmg
Mount, drag to Applications, and configure your API key in Settings.
Choose your format: .AppImage, .deb, or .rpm
Install and run. Requires Python 3.11+.
Configure API key in Settings. First launch may need time for env initialization.
Requires Python 3.11+. The desktop app starts the local daemon automatically. Browser-only UI development requires a manually started daemon.
Natural language commands for anything on your system — from simple queries to complex multi-step automations.
Capability guides
Plain-language guides explain the real workflow, safety boundary, hardware requirements, and known limits for every major interaction mode.
Continuous listening, speech, approvals, and device limits.
02Target resolution, visible actions, and verified outcomes.
03Multiple input paths with independent stop controls.
04On-device signals, temporal checks, and calibration.
05Durable jobs and bounded execution—not unlimited authority.
06Reviewed submissions, integrity checks, and capabilities.
07 · RESEARCH BOUNDARYSynthetic and recorded signal validation, clearly separated from any live brain-control claim.
Hi, I'm Vyom Kulshrestha, a pre-final year Computer Science student at VIT, passionate about building next-generation AI systems and intelligent developer experiences.
I'm building Heliox OS — a futuristic AI-powered operating system interface designed to combine autonomous agents, voice interaction, multimodal intelligence, and system-level automation into a seamless computing experience.
Whether you're a developer, open-source contributor, recruiter, or someone excited about the future of human-computer interaction, I'd love to connect.
An open-source project built solo, growing with every contributor. Be part of what comes next.