Technology / Real-Time Signal Extraction

CYGNUS Lite

Current status: CYGNUS Lite is built and configured, but it is not yet wired into production. The page describes the implemented configuration and its target operating model.

Live conversation doesn't wait. When a person is sitting across from an AI avatar, speaking naturally, reacting in real time, the perception system has to keep up. Every millisecond of processing delay is a millisecond where the avatar is responding to stale information. In real-time interaction, speed isn't a nice-to-have. It's the difference between perception that feels alive and perception that feels disconnected.

CYGNUS Lite is the real-time configuration of the CYGNUS signal extraction engine. It processes the same three channels as CYGNUS Standard (facial, vocal, and postural), but it's been engineered from the ground up for speed. Every decision in CYGNUS Lite's architecture optimizes for one thing: delivering accurate behavioral signals fast enough to keep pace with live human conversation.

This isn't a stripped-down version of CYGNUS Standard. It's a purpose-built instrument for a fundamentally different operating environment.

Listen mode

A guided audio version of this page, adapted for clarity and flow rather than read word for word.

0:00/ 0:00
Loading
Curated narration
Live stack
Camera/Mic→
CYGNUS Lite→
ORACLE RT→
GLUE→
Language Model→
Avatar Response
The Real-Time Problem

Speed is part of the perception quality

CYGNUS Standard processes up to 45 Action Units, the full prosodic feature set, and the complete postural parameter set at maximum resolution. For offline analysis, deep perception sessions, or research, that depth is exactly right. But maximum resolution comes with computational cost, and computational cost means latency.

In a live conversation, latency kills the experience. If the avatar takes 2 seconds to process what it just saw, the person has already moved on to a new thought, a new expression, a new emotional state. The avatar's response will be based on information that's already outdated. Worse, the person will feel the lag. Conversation has a rhythm, and perception that can't keep up with that rhythm breaks it.

CYGNUS Lite solves this by making intelligent tradeoffs between depth and speed. It doesn't process less. It processes smarter.

How Lite Differs from Standard

Same three channels. Different operating philosophy.

01

Signal Selection

CYGNUS Standard extracts all 45 Action Units simultaneously. CYGNUS Lite extracts a priority subset: the Action Units that carry the most behavioral information in conversational contexts.

Research consistently shows that a subset of Action Units accounts for the vast majority of meaningful facial signal variation in conversation. Brow movements (AU1, AU2, AU4), eye region dynamics (AU5, AU6, AU7), and mouth movements (AU12, AU15, AU17, AU20, AU24) carry more conversational signal per unit of computation than the full set. CYGNUS Lite focuses its processing budget on these high-value signals.

This doesn't mean the other Action Units are irrelevant. For deep analysis, subtle signals like AU14 or AU22 can be revealing. But in a live conversation where the avatar needs to respond in sub-second timescales, the priority set captures what matters most for real-time behavioral awareness.

02

Processing Pipeline

CYGNUS Standard runs its extraction pipeline in depth-first mode: every frame gets full processing before results are emitted. CYGNUS Lite runs in speed-first mode: results are emitted as soon as the priority signals are extracted, and additional processing happens in the background if spare compute cycles are available.

This means CYGNUS Lite's output arrives faster and more frequently. Instead of waiting for a complete extraction pass, the avatar receives continuous signal updates at a rate that matches conversational tempo.

03

Vocal Feature Set

The vocal channel in CYGNUS Lite extracts core prosodic features: fundamental frequency, speech rate, volume dynamics, and pause detection. These are the vocal signals most relevant for real-time conversational adaptation. CYGNUS Lite keeps the vocal channel strong by prioritizing the signals that drive live responsiveness first, while deeper secondary markers can remain available for non-live or extended analysis modes.

04

Postural Tracking

CYGNUS Lite tracks the primary postural parameters: head position, torso orientation, and a composite activity index. The design focus is immediate behavioral relevance inside live interaction. Standard mode can add deeper postural granularity, but Lite still preserves the signals that matter most for moment-to-moment adaptation.

What CYGNUS Lite Enables

Signals that arrive in time to matter

01

Avatar Responsiveness

When CYGNUS Lite feeds signals into the avatar pipeline, the avatar can adapt its behavior within the natural flow of conversation. If the person's facial signals shift, the avatar can adjust its tone in its next response. If the person's speech rate drops and pauses lengthen, the avatar can slow down and give more space. If postural engagement increases, the avatar can match that energy.

None of this works if the signal data arrives too late. CYGNUS Lite's sub-second delivery ensures the avatar is always responding to what's happening now, not what happened three seconds ago.

02

Paired with ORACLE RT

CYGNUS Lite provides the signal stream. ORACLE RT provides real-time pattern recognition on that stream. Together, they form a complete real-time perception system: signals are extracted and patterns are detected within the same sub-second window.

This pairing is the core of live avatar perception. CYGNUS Lite tells the system what the face, voice, and body are doing right now. ORACLE RT tells the system what those signals mean in combination. GLUE packages this into context that the language model can use to generate an informed response.

03

Paired with CANON

CANON's personal baseline calibration enhances every CYGNUS Lite extraction. When CANON has accumulated enough data about a person's typical facial dynamics, vocal range, and postural patterns, CYGNUS Lite's signals become personally meaningful.

In live sessions, CANON provides the baseline context that makes CYGNUS Lite's signals interpretable at the individual level. The signals themselves are the same regardless of CANON's status. But the downstream interpretation becomes dramatically more precise with personal calibration.

Three Channels, Optimized

Standard vs Lite (Channel by Channel)

Lite is not the weaker column. It is the live column. Standard maximizes total depth. Lite maximizes signal usefulness per millisecond, so the system can stay inside the rhythm of an active conversation.

Standard: maximum depthLite: sub-second delivery target2s delay already breaks live rhythm
Facial Channel (Priority Set)
Feature
Standard
Lite
Spatial Landmarks
Up to 500
Reduced set, high-signal regions
Action Units
Full set (AU1-45)
Priority set (~15-20 AUs)
Update Rate
Multiple per second
Continuous, speed-optimized
Focus
Complete facial mapping
Conversational signal density
Vocal Channel (Core Set)
Feature
Standard
Lite
Prosodic Features
Full set (7+)
Core set (4-5)
Vocal Quality
Full analysis
Real-time priority tracking
Rhythm Analysis
Full
Conversational rhythm focus
Update Rate
Multiple per second
Continuous, speed-optimized
Postural Channel (Primary Set)
Feature
Standard
Lite
Head Tracking
Full (tilt, nod, rotation)
Full
Shoulder Analysis
Detailed
Primary live profile
Torso Orientation
Full
Primary angles
Activity Index
Full composite
Real-time activity composite
Micro-gestures
Detected
Extended-analysis mode
Use Lite vs Standard

The choice is about context, not quality

The choice isn't about quality. It's about context.

Use CYGNUS Lite when the interaction is live. An avatar is responding in real time. A coaching session is happening synchronously. A video call is active. Any scenario where the person expects natural conversational pacing and where perception delay would degrade the experience.

Use CYGNUS Standard when the content has already been recorded. An uploaded video needs deep analysis. A research dataset is being processed. A session recording is being reviewed after the fact. Any scenario where depth matters more than speed and where the person isn't waiting for a real-time response.

Both produce valid, accurate signal data. The difference is which signals are prioritized and how fast they're delivered. CYGNUS Lite gives you the most important signals immediately. CYGNUS Standard gives you all the signals at maximum depth.

Data Architecture

Same privacy model, narrower live scope

CYGNUS Lite produces the same data classification as CYGNUS Standard: Raw Data. The output is anonymous numerical values with no identity attached. No images are stored. No audio recordings persist. Frames and audio are consumed by the extraction process and discarded immediately.

The only difference is the scope of the output: fewer data points per frame, but at the same quality within those data points. A priority Action Unit extracted by CYGNUS Lite has the same precision and the same measurement methodology as the same Action Unit extracted by CYGNUS Standard.

Open Source Availability

A real-time extraction core developers can build on

CYGNUS Lite is available as an open-source component for developers who want to integrate real-time signal extraction into their own systems. The open-source release includes the core extraction pipeline and can be integrated with custom downstream processing or paired with ORACLE RT for a complete real-time perception stack.

For developers building AI avatars, virtual agents, or interactive systems that need behavioral awareness, CYGNUS Lite provides a foundation that can be deployed independently or as part of the broader OPM pipeline.

Frequently Asked Questions

The practical questions around speed, scope, and deployment

For the signals it extracts, no. Each individual Action Unit, prosodic feature, or postural parameter is measured with the same methodology and the same precision. CYGNUS Lite extracts fewer signals, not worse signals. The priority set is the most informative subset for conversational contexts.

Technical Summary

Operational boundaries in one live frame

Input
Live video + audio feed
Output
Real-time multi-channel signal stream (priority sets)
Facial Data
Priority Action Unit set (~15-20 AUs), reduced landmark set
Vocal Data
Core prosodic features (pitch, rate, volume, pauses)
Postural Data
Primary parameters (head, torso, activity)
Processing Mode
Speed-first, continuous emission
Latency Target
Sub-second within avatar pipeline
Data Storage
None. Frames and audio discarded after extraction.
Data Classification
Raw Data (anonymous, non-personal)
Open Source
Available
Best Paired With
ORACLE RT (real-time pattern recognition)