A mocap camera never sees the actor, only bright dots on black. Optical mocap happens on a stage ringed by 12 to 48 infrared cameras. The performer wears 40 to 60 markers covered in retroreflective tape, which bounce the cameras' infrared light straight back.
A mocap camera never sees the actor, only bright dots on black. Walk a faceless performer through a ring of infrared cameras, cross the rays that pin each marker to a fraction of a millimetre, hand the motion to a four-armed creature, solve a face from helmet-camera dots, and watch an inertial suit drift away.
MocapClearOpened 6 Sept 202614 min to playFree · no sign-up
In 60 seconds
The capture volume
Optical mocap happens on a stage ringed by 12 to 48 infrared cameras. The performer wears 40 to 60 markers covered in retroreflective tape, which bounce the cameras' infrared light straight back. Each camera sees only bright dots on black and sends their 2D centres, about 120 times a second.
Where rays meet
One camera only gives a direction to a dot: a ray. Two or more rays from different angles cross at the marker's 3D position. Centres found to a fraction of a pixel place a still marker to well under a millimetre (0.15 mm in one lab test). When an arm hides a marker from all but one camera, there is a gap.
Dots to a creature
Software names every dot (labelling), fits a skeleton inside them (solving), then copies the joint rotations onto a creature (retargeting). If the creature's legs are longer but its body travels the actor's distance, its planted feet slide. Scaling the travel or locking the feet fixes it.
Faces and fingers
A head-mounted camera watches dots on the face. A facial solver finds the mix of blend shapes (jaw open, smile, brows up) that explains how the dots moved, and those weights drive the creature's face, even remapped to ears. Gloves with bend sensors capture fingers.
Other ways to capture
Inertial suits use about 17 motion sensors and work anywhere, but position comes from adding up acceleration twice, so it drifts unless planted feet reset it. Markerless systems find body keypoints in ordinary video with machine learning, and phones capture faces with 30,000 infrared dots.
On set and cleanup
In performance capture, a tracked virtual camera shows the director the creature live in its world. Afterwards, gaps are filled with smooth curves and jitter is filtered out, but heavy smoothing flattens real motion. Animators then polish, for films shot by shot and for games as loops.
The history
From a galloping horse and a man in a striped suit to creatures that act live on set: 150 years of turning movement into numbers.
A ring of infrared cameras watches a performer covered in shiny dots.
Motion capture (mocap) records how a real person moves, so a computer character can move the same way. The most accurate kind is optical mocap. It happens in a capture volume: an empty stage ringed by special cameras, often 12 to 48 of them.
The performer wears a tight black suit with about 40 to 60 markers: small balls covered in retroreflective tape, the same stuff as on road signs and safety jackets. Each camera has a ring of infrared LEDs around its lens. The tape bounces that light straight back into the lens, so the markers shine brightly while everything else looks dark.
So a mocap camera does not film a picture of the actor. It sees only bright dots on black, finds the centre of each dot, and sends those 2D positions to the computer, about 120 times a second. Look at the four camera views: each camera sees a different pattern, and some dots are missing because the body is in the way.
One camera gives a direction. Two or more give a point in 3D.
A single camera cannot tell how far away a dot is. It only knows the direction: the marker is somewhere along a line, a ray, from the lens through that dot. Hold one eye shut and try to touch two pencil tips together: it is surprisingly hard.
A second camera, somewhere else, gives a second ray. Where the two rays cross is the marker. This is triangulation, the same idea as your two eyes judging distance. With more cameras the rays pin the point down better, and it helps when the cameras look from very different angles. Before a shoot, the team waves a calibration wand through the volume so the computer knows exactly where each camera is.
Because each dot's centre is found to a fraction of a pixel, a good optical system places a still marker to well under a millimetre. A lab test of one commercial system measured an average error of 0.15 mm. The big enemy is occlusion: when an arm or another actor blocks a marker, fewer cameras see it. Under two, it cannot be placed at all, and a gap appears in the data.
Name the dots, fit a skeleton, then hand the motion to someone else’s body.
Triangulation gives a cloud of 3D dots, 120 times a second, but the computer does not yet know which dot is which. Step one is labelling: each dot gets a name like LKNE (left knee) or RWRA (right wrist). The software uses a template of the suit and follows each dot from frame to frame. When two dots pass close together, it can mix them up, a marker swap, which a person then fixes.
Step two is solving. The software fits a skeleton (a set of bones and joints sized to the actor) inside the labelled dots, frame by frame. The result is not dots any more but joint rotations: how much the hip, knee and elbow bend.
Step three is retargeting: copying those rotations onto a character with a different body. The character needs its own skeleton and rig (see Anim3DClear for rigging). If its legs are longer, each copied step is longer too. If the body still travels as far as the actor did, the planted foot must slip along the floor: foot slide. The fix is to scale the travel to the new leg length, and to lock the feet to the floor with IK (inverse kinematics).
A camera on a boom watches dots on the face; gloves read the fingers.
Bodies are only half a performance. For the face, the actor wears a head-mounted camera (HMC): a small camera on a boom fixed to a helmet, pointing back at the face. Because it moves with the head, the face stays in the same spot in its picture, so only expressions make the dots move. Dots are painted or stuck on the skin: dozens, up to about 150.
The facial solver works backwards. It knows how the dots move for each basic expression, called a blend shape: jaw open, smile, brows up, blink. It finds the mix of blend shapes that best explains what the camera saw. Those numbers, 0 to 1 for each shape, then drive the creature’s own blend shapes. Artists can remap them: a creature with big ears could lift its ears when the actor lifts their brows.
Fingers are small and hide each other, so they are hard for body cameras. Performers often wear gloves with bend sensors along each finger, or tiny markers, or animators add fingers by hand.
Motion sensors in a suit, AI that watches ordinary video, and your phone.
Optical markers are the most accurate, but you need a stage full of cameras. There are other ways.
An inertial suit has about 17 small motion sensors (IMUs), one on each body segment. Each has a gyroscope, an accelerometer and a compass (magnetometer), like the ones in your phone. They measure how each bone turns, very well, and they work anywhere, even outdoors. But to know where you are, the suit must add up acceleration twice, and tiny errors grow fast: this is drift. Suits fight it by knowing that a planted foot is not moving. Steel in a floor can also upset the compass.
Markerless capture uses ordinary video and machine learning (pose estimation) to find keypoints like elbows and knees, with no suit at all. OpenPose (2017) did this in real time; Google’s BlazePose (2020) finds 33 body points on a phone. It is cheap and quick, but less precise, and it guesses what it cannot see.
Your phone can capture faces: a depth camera projects over 30,000 invisible dots, and the software gives 52 expression numbers, enough to drive a cartoon face live.
The director sees the creature live, then artists clean and polish the data.
On a film, body, face and voice are often recorded together: performance capture. The director does not want to imagine the creature, so the stage streams the solved skeleton into a game engine in real time. A camera operator carries a virtual camera: a screen with markers on it. The system tracks it like a marker, and the screen shows the creature in its virtual world, framed from where the operator stands. The next step, filming real sets with virtual backgrounds on LED walls, is VirtualProdClear.
The live preview is rough. Afterwards the data is cleaned. Gaps where a marker was hidden are filled with a smooth curve. Jitter, the small shake of every measurement, is removed with a smoothing filter. Filter too hard and the motion goes soft: a quick hand flick or the top of a jump gets cut down.
Then animators polish: fixing feet, adding fingers, pushing a pose to be more readable, or blending in keyframes (see Anim3DClear). Games record short moves (walk, turn, jump) that loop and blend live at 30 to 60 frames a second, so they must join seamlessly. Films capture whole scenes, and each shot gets hours of care. Creatures are then finished with muscles, skin and lighting (see VFXClear).
The 2D positions of bright marker dots. Infrared light bounces back from the markers, so each camera sees bright dots on black and sends the centre of each dot.
Why are the markers covered in retroreflective tape?
It bounces the cameras’ infrared light straight back into the lens. Retroreflective tape sends light back the way it came, so the markers shine brightly for the camera that lit them.
Why do studios use many cameras all around the stage?
So every marker is seen by several cameras even when the body hides it from some. Arms, legs and the body block markers from some cameras. With cameras all round, others still see them.
What can a single camera tell you about a marker?
Only the direction to it (a ray). A dot in one picture could be anywhere along a line from the lens. You need a second view to know where along the line.
A marker is seen by only one camera for a few frames. What happens?
A gap appears in its 3D track. With fewer than two rays, the point cannot be triangulated, so those frames are a gap that must be filled later.
Which pair of cameras places a marker most accurately?
Two cameras at right angles to each other. Rays that cross at a wide angle pin the point down in every direction. Nearly parallel rays leave it fuzzy along their length.
What does “labelling” mean in mocap?
Deciding which marker each 3D dot is, in every frame. Triangulated dots have no names. Labelling matches each one to a marker on the suit template, frame by frame.
What comes out of solving?
A skeleton’s joint positions and rotations. The solver fits bones inside the markers. Animators work with the joint rotations, not the dots.
A creature has legs 1.5 times as long as the actor. Its body travel is copied as-is. What goes wrong?
Its feet slide on the ground. Each step is 1.5 times longer but the body moves the actor’s distance, so the planted foot must slip. Scale the travel or lock the feet.
Why is the face camera mounted on the actor’s helmet?
So it moves with the head and only expressions move the dots. With the camera fixed to the head, head turns do not move the face in the picture, so every dot movement is an expression.
What does the facial solver output?
How much of each blend shape (0 to 1) is in the expression. It finds the mix of blend shapes that best explains the dot movements, and those weights drive the creature’s face.
Why are fingers often captured with gloves?
Fingers are small and hide each other from body cameras. Fingers are tiny and constantly block one another, so bend sensors in gloves (or animators) fill in the detail.
What does an inertial suit use instead of cameras?
Small motion sensors (IMUs) on each body part. Each IMU has a gyroscope, accelerometer and magnetometer that measure how that body part moves and turns.
Why does an inertial suit’s position drift?
Tiny acceleration errors are added up twice, so they keep growing. Position comes from integrating acceleration twice. A tiny constant error grows with time squared unless foot contacts reset it.
What is the big advantage of markerless video capture?
No suit or markers, just cameras and software. Machine learning finds the pose in ordinary video. It is quick and cheap, but less precise than markers.
What does the virtual camera’s screen show?
The creature in its digital world, from where the operator stands. The virtual camera is tracked, and a game engine renders the creature and set from its position, live.
How are gaps from hidden markers usually fixed?
A smooth curve is drawn across the missing frames. Cleanup software fills a gap with a smooth curve that joins the good data on either side.
What goes wrong if you smooth mocap data too much?
Quick moves get softened and peaks are cut down. A heavy low-pass filter removes fast changes too, so snaps, hits and the top of a jump lose their sharpness.
Words worth knowing
Motion capture
Recording a real performer's movement so a digital character can move the same way.
Retroreflective marker
A small ball with tape that bounces light straight back to the camera that lit it.
Triangulation
Finding a point in 3D from where rays from two or more cameras cross.
Occlusion
When a body part or object hides a marker from a camera.
Solving
Fitting a skeleton inside the labelled markers to get joint rotations.
Retargeting
Copying motion from one skeleton onto a body with different proportions.
Blend shape
A saved face expression mixed in by a weight from 0 to 1.
IMU
A tiny gyroscope, accelerometer and compass that measures how a body part moves.
Performance capture
Capturing body, face and voice together so the whole acting performance is kept.
Fork it. Teach with it.
This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/mocapclear”.