VisionClaw

VisionClaw is a useful public-source, terms-constrained wearable-agent prototype and architecture stress test. Its current iOS main separates LiveKit realtime perception from a hosted action agent; its README, saved X post, and Android client still describe an older direct Gemini Live + optional OpenClaw design. Keep the signal and the history, but do not treat the project as a production-ready install.

Decision

Retain as an architecture reference; do not install or recommend as a default runtime. Bind its durable patterns and counterexamples to Realtime Voice Agent Workflow. Reconsider adoption only after a pinned revision has a usable license grant, explicit human tool confirmation, least-privilege session identity, durable deferred jobs, pinned dependencies/models, platform-specific proof, and representative tests.

The saved image is a screenshot of promotional README material with composed lifestyle imagery. It supports what the source claimed; it is not proof that the reviewed revision ran on glasses.

Frozen source

Field Reviewed value
Repository revision Intent-Lab/VisionClaw@675a0bf2595773d09e4a14c8d8a03dba94945044, 2026-08-06
GitHub snapshot 2,500 stars, 488 forks, 40 open issues, 167 source files
Release identity No GitHub releases; v1.0 points to older commit b1a0de8182c48e21166a5f8b6c1a23440c72797f
License GitHub NOASSERTION / Other; repository LICENSE only invokes Meta Wearables Developer Terms and AUP
Supply-chain check Current gateway lockfile: 81 production dependencies and 0 known npm audit --omit=dev advisories; not a security proof
Proof completed here Exact source, commit history, mobile/server paths, source image, official Gemini/LiveKit/Meta sources, and file hashes reviewed; no hardware run

The X post's “100% Open Source” language is too broad. The code is publicly readable, but the repository contains no general permissive license and its wearable dependency is governed by Meta terms. Source: exact repository LICENSE and GitHub metadata, 2026-08-11

Current architecture, not the README diagram

The iOS client now publishes microphone and camera tracks into a fresh LiveKit room and subscribes to agent audio. The worker owns the Gemini or OpenAI realtime provider key and can delegate a fast search or a longer task. Longer work crosses the hosted gateway into one shared Anthropic agent/environment with per-user vaults and sessions. It waits for a bounded interval, then pushes a late result or parks it for the next call. A pinned/current camera frame can travel with the delegated task so the action model sees the source pixels instead of only a voice model's summary. Source: exact LiveKitSession.swift, agent/main.py, and gateway/src at the reviewed commit

The gateway deliberately retains an OpenAI-compatible/OpenClaw-compatible protocol boundary. That does not make the current iOS action plane OpenClaw. The older README and current Android source still use direct Gemini Live and optional OpenClaw; the current iOS source uses LiveKit plus Anthropic Managed Agents. The README's “56+ skills” claim is therefore historical/platform-specific, not a current cross-platform capability claim.

What the source teaches the harness

  • Separate perception from action. The realtime loop needs interruption, audio, room, and camera semantics; the action loop needs authority, durable execution, tools, and receipts. They can change independently.
  • Name every processor. Camera/mic capture, LiveKit, realtime model, search, action model, gateway, vault, app MCPs, storage, and analytics are separate trust and cost boundaries.
  • Bind dispatch to session semantics. VisionClaw moved to a unique room per call because its agent dispatch fires when a room is created; reusing a room produced dead redials.
  • Preserve visual reference. “This” must resolve to an explicit current or pinned frame. Attach the authoritative frame when a downstream task depends on dense visual detail.
  • Do not block conversation on long work. Use a bounded synchronous wait, progress speech, durable job identity, and at-most-once late-result delivery now or on the next call.
  • Budget history. A bounded recent-task briefing is preferable to injecting unbounded action history into a latency- and cost-sensitive realtime context.
  • Expose liveness. Distinguish connected room, waiting for worker, model starting, listening, thinking, speaking, degraded audio, and voice-only fallback.
  • Measure memory before adding it. This source removed memory writes from the critical path after observing 5–10 seconds of overhead.

These are requirements in Realtime Voice Agent Workflow, not a reason to copy this repository wholesale.

Blocking defects and risks

Tool approval is not approval

Current gateway/src/turn.ts automatically answers every Anthropic requires_action event with user.tool_confirmation: allow. The source itself says destructive tools should later receive spoken confirmation, while gateway/src/provision.ts currently gives connected MCP toolsets always_allow. If a future tool switches to always_ask, the same handler still approves it automatically. This defeats the policy boundary.

Kevin's harness must never let a model, transport, or recovery loop answer its own confirmation. A spoken confirmation state must name the action, target, material parameters, and consequence; record the human answer; expire on silence/disconnect; reject stale answers; and produce a receipt.

Admission gaps

  • The action environment has unrestricted networking; one shared environment/agent serves per-user sessions and vaults.
  • A gateway service token plus X-User-Id can impersonate users. Identity must be bound server-side, rotated, and audited.
  • Room tokens are short-lived, but current grants are broader than source-specific mic/camera publish permissions.
  • The Fly-volume JSON store lacks application-layer encryption and is coupled to cloud-resource reconciliation.
  • Deferred work drains in process, so a restart may drop an in-flight job.
  • Python dependencies use moving lower bounds, preview model identifiers are not locked to the repo commit, and a failed model prefetch does not fail the image build.
  • The repository has no CI and only skeletal default mobile tests; no representative hardware, reconnection, voice, confirmation, or cross-platform suite was found.
  • Android still permits cleartext traffic, persists configuration in SharedPreferences, uses debug signing for release, and disables minification. It is not parity proof.
  • Meta DAT is a developer-preview dependency. Meta's SDK repositories state analytics and crash reporting are enabled by default unless disabled, and Meta terms/integration review apply.

Primary-source constraints

Google documents Gemini Live JPEG visual input at no more than one frame per second, raw 16 kHz PCM audio input, and 24 kHz PCM output. Google recommends short-lived ephemeral tokens for production client-to-server use, immediate local playback-buffer clearing on interruption, and explicit resumption/compression/GoAway handling for longer sessions. Accumulated context and proactive listening also affect cost. Source: Gemini Live API, accessed 2026-08-11

LiveKit access tokens can be short-lived and room-scoped, with granular publish permissions including allowed track sources. Explicit agent dispatch attached to a participant token applies when the room is created, which supports the source's fresh-room-per-call correction. Source: LiveKit token authentication and agent dispatch, accessed 2026-08-11

Promotion checklist

Before this can become more than a reference:

  1. obtain an actual software license and complete Meta terms/data review;
  2. choose and pin one platform path, provider/model set, SDK set, and deployment revision;
  3. remove automatic confirmation and prove spoken approve/deny/timeout/reconnect behavior;
  4. scope room media grants and bind all user identity server-side;
  5. move deferred tasks to durable, idempotent job execution with single-delivery receipts;
  6. encrypt/minimize persisted data and prove OAuth revoke/delete/recovery;
  7. add deterministic state/protocol tests plus representative device, network, voice, visual, cost, and adversarial evaluations;
  8. verify analytics/crash defaults, visible recording state, bystander consent, and camera/mic deletion.

Timeline

  • 2026-08-11 | Replayed the source against exact current main and official Gemini, LiveKit, and Meta evidence. Corrected the stale README/X architecture: current iOS uses LiveKit plus a hosted Anthropic action plane, while Android remains on the older direct Gemini/OpenClaw path. Promoted perception/action separation, frame pinning, fresh-call rooms, bounded late-result delivery, liveness states, and context budgeting; recorded the automatic-confirmation defect, license/terms limits, platform divergence, identity/persistence risks, and missing proof. Source: .brain/artifacts/x/2037892504683716650/visionclaw-review/source-audit.md
  • 2026-07-04 | Created from the X bookmark and repository README. That snapshot described direct Gemini Live plus optional OpenClaw and is retained as historical source state, not current iOS truth. Source: X/@ihtesham2005, 2026-03-28