The full three-channel configuration. Facial, vocal, and postural extraction running simultaneously at maximum resolution.
Built for scenarios where depth of observation matters more than speed.
Deployed in ANIMA and Silent Oculus.
A guided audio version of this page, adapted for clarity and flow rather than read word for word.
Every interaction begins with a signal. A shift in the brow. A change in vocal tone. A subtle adjustment in posture. These signals happen constantly, involuntarily, and faster than any human can consciously track. CYGNUS is the engine that captures them all.
CYGNUS is the signal extraction model of the OPM pipeline. It operates at the very front of the perception chain, converting raw sensory input from three distinct channels into structured, anonymous numerical data. Instead of interpreting, classifying, or assigning meaning to what it sees, CYGNUS simply extracts. Pure measurement. Pure signal.
Think of CYGNUS as the retina of the system. The retina does not understand what it sees. It converts light into electrical signals and sends them deeper into the brain. CYGNUS does the same with human expression: it converts observable behavior into structured data and passes it forward.
CYGNUS does not interpret. It extracts observable human expression across facial, vocal, and postural channels and turns raw media into anonymous numerical signal streams.
Human communication is inherently multi-channel. What someone says with their face may not match what they say with their voice, and neither may match what their body is doing. CYGNUS captures all three channels simultaneously, creating a complete signal picture that no single human observer could maintain across an entire conversation.
The human face is the most expressive surface on the body. It contains over 40 independently controllable muscle groups, each producing distinct visible movements. These movements are catalogued in behavioral science as Action Units, a standardized system that maps every observable facial movement to a numerical code.
CYGNUS maps the face through up to 500 spatial landmarks, capturing the geometry and micro-movements of every facial region with sub-millimeter precision. These landmarks cover the forehead, brow ridge, eye contours, nasal bridge, cheekbones, mouth corners, jawline, and chin, creating a high-resolution spatial mesh that tracks even the smallest muscular shifts in real time.
From this spatial map, CYGNUS derives up to 45 Action Units as continuous intensity values. Each Action Unit sits on a scale between 0 and 1, where 0 means no activation and 1 means full activation. This is not a photograph or an image. It is a set of numbers describing how the face moves at any given moment.
This is a fraction of the full set. CYGNUS operates across all 45 Action Units simultaneously, updating every value multiple times per second. The result is a continuous, high-resolution stream of facial signal data that describes the face in purely numerical terms.
No images are stored. No photographs are taken. The face enters the system as pixels and leaves as numbers. The pixels are discarded. The numbers move forward.
The inner portion of the eyebrows moves upward. Often associated with attention, surprise, or concern, but CYGNUS does not assign any of these meanings. It records: "AU1 = 0.64."
The eyebrows pull together and downward. This is one of the most frequently detected Action Units in conversational settings. CYGNUS records its intensity and duration without interpreting why it occurs.
The cheeks push upward, causing the skin below the eyes to gather. Its presence or absence alongside other AUs carries meaning, but that meaning lives downstream. CYGNUS records the number.
The corners of the mouth pull upward. CYGNUS detects its intensity, duration, and symmetry. Whether it occurs with or without AU6 is data that CYGNUS records. What that combination means is not CYGNUS's concern.
Subtle chin raises and lip pressing often happen quickly and are frequently missed by human observers. CYGNUS captures them consistently, regardless of how subtle or brief they are.
The voice carries more information than words. The same sentence spoken with a rising tone, a falling tone, a steady tone, or a breaking tone communicates fundamentally different things. CYGNUS captures the non-verbal properties of the voice, everything except the words themselves.
Fundamental frequency, pitch variability, speech rate, volume dynamics, pause architecture, vocal quality markers, rhythm, and cadence are all extracted in real time and converted to numerical values. The audio stream is consumed by the extraction process, and what survives is pure measurement. No voice recording exists after processing.
The body communicates constantly, and most of it happens below conscious awareness. Posture shifts, weight redistributions, arm positions, head angles, shoulder tension. These are signals that even trained observers struggle to track consistently because they change gradually and subtly.
CYGNUS captures postural data as a set of numerical values representing body position and movement: head position, shoulder line, torso orientation, gestural activity, and openness or closedness on a continuous scale. The video frames are processed and immediately discarded. What remains is a stream of numbers describing body position and how it changes over time. No images of the body are stored.
The output of CYGNUS is a structured, multi-channel signal stream. At any given moment, this stream contains facial Action Unit intensities, prosodic feature dimensions, and postural parameters updating multiple times per second.
Understanding CYGNUS means understanding its boundaries. These are not limitations. They are design decisions.
The full three-channel configuration. Facial, vocal, and postural extraction running simultaneously at maximum resolution.
Built for scenarios where depth of observation matters more than speed.
Deployed in ANIMA and Silent Oculus.
The real-time configuration. All three channels remain active, but the feature set is curated for maximum relevance at minimum latency.
Engineered for live, face-to-face perception where the avatar must react within the natural flow of conversation.
Deployed in Anima Connect and WhatsAnima.
The audio-only configuration. No camera required.
Extracts the full vocal signal set from audio input alone and is a complete perception instrument in its own right.
Deployed in WhatsAnima voice messages and Anima Connect audio-only mode.
No. CYGNUS detects that a face is present and measures how it moves. It does not compare faces to databases, does not store facial images, and cannot determine who a person is.
No. Video frames and audio waveforms are processed in real time and discarded immediately after signal extraction. The output is numerical data only.
Raw Data produced by CYGNUS is anonymous numerical signal data. When linked to an identifiable individual or session, it becomes Attributed Data and is treated as personal data.
Yes. CYGNUS ECHO operates on audio input only and delivers the full vocal signal set without any camera input.
CYGNUS Lite is specifically optimized for real-time processing in live conversational scenarios. CYGNUS Standard prioritizes depth over speed.
Emotion detection systems output labels like "happy" or "angry." CYGNUS outputs numbers: "AU4 = 0.73, AU12 = 0.81, pitch = 214Hz." Numbers describe what is happening. Labels claim to know why.
CYGNUS is the foundation of the OPM pipeline. For how CYGNUS data is analyzed across channels, the next step is ORACLE.