Technology / Signal extraction model

CYGNUS

Listen Mode

A guided audio version of this page, adapted for clarity and flow rather than read word for word.

0:00/ 0:00
Loading
Curated narration

Every interaction begins with a signal. A shift in the brow. A change in vocal tone. A subtle adjustment in posture. These signals happen constantly, involuntarily, and faster than any human can consciously track. CYGNUS is the engine that captures them all.

CYGNUS is the signal extraction model of the OPM pipeline. It operates at the very front of the perception chain, converting raw sensory input from three distinct channels into structured, anonymous numerical data. Instead of interpreting, classifying, or assigning meaning to what it sees, CYGNUS simply extracts. Pure measurement. Pure signal.

Think of CYGNUS as the retina of the system. The retina does not understand what it sees. It converts light into electrical signals and sends them deeper into the brain. CYGNUS does the same with human expression: it converts observable behavior into structured data and passes it forward.

Core Thesis

CYGNUS does not interpret. It extracts observable human expression across facial, vocal, and postural channels and turns raw media into anonymous numerical signal streams.

3
Channels
45
Action Units
500
Spatial landmarks
Pipeline Position
Input (Video / Audio)→CYGNUS→Raw Signal Data→ORACLE→Pattern Analysis
Three Channels, One Engine

Human communication is inherently multi-channel.

Human communication is inherently multi-channel. What someone says with their face may not match what they say with their voice, and neither may match what their body is doing. CYGNUS captures all three channels simultaneously, creating a complete signal picture that no single human observer could maintain across an entire conversation.

Facial Channel

The face enters as pixels and leaves as numbers.

The human face is the most expressive surface on the body. It contains over 40 independently controllable muscle groups, each producing distinct visible movements. These movements are catalogued in behavioral science as Action Units, a standardized system that maps every observable facial movement to a numerical code.

CYGNUS maps the face through up to 500 spatial landmarks, capturing the geometry and micro-movements of every facial region with sub-millimeter precision. These landmarks cover the forehead, brow ridge, eye contours, nasal bridge, cheekbones, mouth corners, jawline, and chin, creating a high-resolution spatial mesh that tracks even the smallest muscular shifts in real time.

From this spatial map, CYGNUS derives up to 45 Action Units as continuous intensity values. Each Action Unit sits on a scale between 0 and 1, where 0 means no activation and 1 means full activation. This is not a photograph or an image. It is a set of numbers describing how the face moves at any given moment.

This is a fraction of the full set. CYGNUS operates across all 45 Action Units simultaneously, updating every value multiple times per second. The result is a continuous, high-resolution stream of facial signal data that describes the face in purely numerical terms.

No images are stored. No photographs are taken. The face enters the system as pixels and leaves as numbers. The pixels are discarded. The numbers move forward.

Example Signals

Action Units CYGNUS captures

AU1 (Inner Brow Raise)

The inner portion of the eyebrows moves upward. Often associated with attention, surprise, or concern, but CYGNUS does not assign any of these meanings. It records: "AU1 = 0.64."

AU4 (Brow Lowerer)

The eyebrows pull together and downward. This is one of the most frequently detected Action Units in conversational settings. CYGNUS records its intensity and duration without interpreting why it occurs.

AU6 (Cheek Raiser)

The cheeks push upward, causing the skin below the eyes to gather. Its presence or absence alongside other AUs carries meaning, but that meaning lives downstream. CYGNUS records the number.

AU12 (Lip Corner Puller)

The corners of the mouth pull upward. CYGNUS detects its intensity, duration, and symmetry. Whether it occurs with or without AU6 is data that CYGNUS records. What that combination means is not CYGNUS's concern.

AU17 / AU24

Subtle chin raises and lip pressing often happen quickly and are frequently missed by human observers. CYGNUS captures them consistently, regardless of how subtle or brief they are.

Vocal Channel

The voice carries more information than words.

The voice carries more information than words. The same sentence spoken with a rising tone, a falling tone, a steady tone, or a breaking tone communicates fundamentally different things. CYGNUS captures the non-verbal properties of the voice, everything except the words themselves.

Fundamental frequency, pitch variability, speech rate, volume dynamics, pause architecture, vocal quality markers, rhythm, and cadence are all extracted in real time and converted to numerical values. The audio stream is consumed by the extraction process, and what survives is pure measurement. No voice recording exists after processing.

Prosodic Feature Set

CYGNUS listens for signal, not language.

  • Fundamental Frequency (Pitch): measured in Hertz and tracked as a trajectory over time.
  • Pitch Variability: how much pitch moves within a phrase, a sentence, or a minute.
  • Speech Rate: words per minute, syllables per second, including acceleration and deceleration across a session.
  • Volume Dynamics: sudden drops, gradual increases, sustained whispers, and spikes measured as signal shape rather than label.
  • Pause Architecture: where pauses occur, how long they last, and whether they become more frequent.
  • Vocal Quality Markers: breathiness, roughness, strain, tremor.
  • Rhythm and Cadence: the temporal pattern of speech as it evolves over time.
Postural Channel

The body communicates constantly.

The body communicates constantly, and most of it happens below conscious awareness. Posture shifts, weight redistributions, arm positions, head angles, shoulder tension. These are signals that even trained observers struggle to track consistently because they change gradually and subtly.

CYGNUS captures postural data as a set of numerical values representing body position and movement: head position, shoulder line, torso orientation, gestural activity, and openness or closedness on a continuous scale. The video frames are processed and immediately discarded. What remains is a stream of numbers describing body position and how it changes over time. No images of the body are stored.

Postural Parameters

Continuous body position as data.

  • Head Position: tilt angle, nod angle, rotation angle, measured continuously in degrees.
  • Shoulder Line: asymmetry, elevation, forward roll, and slow changes over the duration of a session.
  • Torso Orientation: forward lean, backward lean, lateral lean, measured against baseline position.
  • Gestural Activity: an activity index describing how much upper-body movement is present.
  • Openness / Closedness: a composite spectrum derived from multiple postural indicators.
Output Stream

What CYGNUS produces

The output of CYGNUS is a structured, multi-channel signal stream. At any given moment, this stream contains facial Action Unit intensities, prosodic feature dimensions, and postural parameters updating multiple times per second.

Channel
Data Points
Example Values
Facial
Up to 45 Action Unit intensities
AU4: 0.73 · AU6: 0.12 · AU12: 0.81 · AU15: 0.04
Vocal
7+ prosodic feature dimensions
Pitch: 214Hz · Rate: 142wpm · Volume: 0.67 · Pause: 1.3s
Postural
5+ body position parameters
Head Tilt: -3.2° · Lean: 8.1° forward · Activity: 0.34
This data is Raw Data as defined in our Privacy Policy and Data Processing Agreement. It is inherently anonymous: a set of numbers with no name, no session identifier, and no timestamp linking it to anyone. Attribution happens later, in a separate process, and only when necessary.
Boundaries

What CYGNUS doesn't do

Understanding CYGNUS means understanding its boundaries. These are not limitations. They are design decisions.

CYGNUS does not perform facial recognition or match faces to databases. It knows a face is present and measures how it moves. Who that face belongs to is outside its scope entirely.
Every frame of video is consumed by the extraction process and discarded. Every audio waveform is analyzed and discarded. After processing, only numbers remain.
There are no emotion labels anywhere in CYGNUS output. The output is numerical values: AU4 = 0.73, Pitch = 214Hz. That is the language CYGNUS speaks.
CYGNUS only describes what is happening right now. Predicting what comes next or analyzing what happened before is not part of its function.
CYGNUS Configurations

Three specialized instruments.

CYGNUS Standard

The full three-channel configuration. Facial, vocal, and postural extraction running simultaneously at maximum resolution.

Built for scenarios where depth of observation matters more than speed.

Deployed in ANIMA and Silent Oculus.

CYGNUS Lite

The real-time configuration. All three channels remain active, but the feature set is curated for maximum relevance at minimum latency.

Engineered for live, face-to-face perception where the avatar must react within the natural flow of conversation.

Deployed in Anima Connect and WhatsAnima.

CYGNUS ECHO

The audio-only configuration. No camera required.

Extracts the full vocal signal set from audio input alone and is a complete perception instrument in its own right.

Deployed in WhatsAnima voice messages and Anima Connect audio-only mode.

Configuration Comparison

Standard, Lite, and ECHO

Feature
Standard
Lite
ECHO
Facial Extraction
Full (AU1-45)
Priority Set
Inactive
Vocal Extraction
Full
Core Set
Full
Postural Extraction
Full
Primary Set
Inactive
Real-Time Optimized
No (depth priority)
Yes (speed priority)
Yes
Camera Required
Yes
Yes
No
Primary Context
Deep analysis, research
Live conversations
Voice-only interaction
Frequently Asked Questions

Common CYGNUS questions

Does CYGNUS use facial recognition?+

No. CYGNUS detects that a face is present and measures how it moves. It does not compare faces to databases, does not store facial images, and cannot determine who a person is.

Does CYGNUS record video or audio?+

No. Video frames and audio waveforms are processed in real time and discarded immediately after signal extraction. The output is numerical data only.

Is CYGNUS data personal data under GDPR?+

Raw Data produced by CYGNUS is anonymous numerical signal data. When linked to an identifiable individual or session, it becomes Attributed Data and is treated as personal data.

Can I use CYGNUS without a camera?+

Yes. CYGNUS ECHO operates on audio input only and delivers the full vocal signal set without any camera input.

Does CYGNUS work in real time?+

CYGNUS Lite is specifically optimized for real-time processing in live conversational scenarios. CYGNUS Standard prioritizes depth over speed.

How is CYGNUS different from emotion detection systems?+

Emotion detection systems output labels like "happy" or "angry." CYGNUS outputs numbers: "AU4 = 0.73, AU12 = 0.81, pitch = 214Hz." Numbers describe what is happening. Labels claim to know why.

Technical Summary

CYGNUS at a glance

Input
Video feed (Standard, Lite) or Audio feed (ECHO)
Output
Multi-channel numerical signal stream
Facial Data Points
Up to 45 Action Units, mapped via up to 500 spatial landmarks
Vocal Data Points
7+ prosodic feature dimensions
Postural Data Points
5+ body position parameters
Processing Mode
Real-time, per-frame, ephemeral
Data Storage
None. Frames and audio discarded after extraction
Data Classification
Raw Data (anonymous, non-personal, property of EXIDEUS LLC)
Configurations
Standard, Lite, ECHO
Next Model

CYGNUS is the foundation of the OPM pipeline. For how CYGNUS data is analyzed across channels, the next step is ORACLE.