How FUT HyperVision works

· 12 min read

The build-up to a conceded goal in one of my own matches: the switch that would have cut the line, the missed instruction in red, the next cut, and the open teammate on the attack before it.

Ultimate Team is the most competitive mode in EA SPORTS FC. Matches are fast, opponents are good, and games are decided by details: a switch made half a second late, a passing lane left open, a teammate nobody saw. I have played it for years, and I built FUT Evolution, a companion site and app for its players, so I spend a lot of time around people trying to get better at it.

Under almost every match-analysis video on YouTube, the same question comes back: could AI watch a match and break it down the way a coach would, from where the players stand and where the ball goes, and turn that into advice for the next game?

What a coach does on a paused frame is not magic. It is perception, geometry and a handful of football rules, and each of those can be built and measured. I took the problem apart that way, one piece at a time, and named the system that came out of it FUT HyperVision. The name is my own, not an EA feature. This post shows how each piece is built and what it is measured against. Whether the coaching calls agree with an expert is the next measurement, not a result yet.

What it does

FUT HyperVision takes a recorded FC match and returns the same video with tactical annotations rendered into it: which defender should have been selected to cut a passing lane, where he could reach it, which teammate was open, and whether the player followed the instruction. The annotations follow the visual language EA uses in the game, so they read as part of the broadcast rather than as a debug overlay.

Under the hood it is a perception and reasoning pipeline over 1080p, 60 fps video. It detects and tracks the players, the officials and a ball about 10 pixels wide, registers each frame to pitch coordinates and grades the fit, compensates for camera motion, reconstructs passes and ownership, and runs a rule-based planner over the whole match before a single pixel is drawn.

Three constraints shape every stage. The footage is rendered, not filmed, so models trained on real football transfer poorly. There is no ground truth for most layers, so every evaluation has to be built. And the output is advice, so a confident wrong answer is worse than no answer.

Results at a glance
Each figure states what it is measured against.
LayerResultMeasured against
Ball detection99.6% recall, 0 wrong objects, 0.50 px median526 hand-audited frames, 9 held-out windows
Player detectionmAP50 0.872, precision 0.934, recall 0.774200 held-out auto-labelled images
Camera-motion compensation1.9 px median at 45 framesobserved pitch markings, 28 frame pairs
Pitch registrationGraded per frame: 56.8% good, 11.7% coarse, 31.5% rejectSelf-consistency only, see limitations
Accuracy figures come from labels the model under test did not produce. The one exception, player detection, is marked as agreement with the auto-labeller.

The pipeline

From a match video to a coaching note
A match video in, a coached match video out. Perception, facts, decisions and drawing are separate steps.
Match video
60 fps, 1080p
Perception, every frame
Players
YOLO26s, ByteTrack, kit split
Ball
YOLO11s at 1280 px, continuity
Pitch
32 keypoints, RANSAC, graded
Facts, decisions, drawing
Reconstruction
identities, teams, owner, passes
Tactical plan
switch, cut, contain
Render
chevrons, arrows, traces
Evaluation gate
every render, before delivery
Three perception models, one reconstruction, one plan, one renderer. The plan is computed once for the whole match; any window of it can be rendered, and every render is checked against the plan before delivery.

Stage 1: player detection

A YOLO26s detector with three classes (player, goalkeeper, referee), run at 960 px and tracked with ByteTrack. The tracker is reset on hard cuts, so a replay does not inherit the identities of the live play it interrupted. Officials and goalkeepers are detected explicitly so the planner can reason about outfield players only.

Current detection on an FC frame: dark team boxes in green, light team boxes in orange, each labelled with its shirt brightness, and the referee boxed in red as dropped.
The detector on a live frame. Green and orange are the two teams, each labelled with the shirt brightness that decided it. The referee is found, then dropped.

On a held-out split of 200 images the model reaches mAP50 0.872 overall (0.955 players, 0.927 referees, 0.735 goalkeepers), with precision 0.934 and recall 0.774. That split comes from our own auto-labeller, so these figures measure agreement with it, not accuracy against human labels.

Officials have to be found, not just ignored: a referee standing in a passing lane looks like one more player, and the switch rule must never name him. The referee class covers the assistant referees on the touchline too. Its training labels come from an auto-labeller that runs a generic person detector, assigns the class from kit-colour rules, and keeps a person only when the band under its feet is at least 30% pitch-coloured pixels, which removes spectators behind the advertising boards.

Ten seconds at normal speed from a diagnostic render: the referee and an assistant referee detected and tracked, with the same grass gate applied under their feet.

Teams are split by brightness, not hue: a box around a running player is mostly grass, and hue measures the grass. The classifier samples the upper chest (22 to 48% of the box height), masks out grass pixels and keeps the median HSV value. The values separate into two clusters with a one-dimensional 2-means. The user's team is the cluster under EA's own marker for the controlled player. For the match shown, the match configuration sets the team instead, along with the attacking direction of each half. Each track votes over all its frames and needs 70% agreement to get a team.

Twelve seconds at normal speed. Each outlined patch is the upper chest the brightness is measured on; the pill is its median HSV value. Dark pills are one team, light pills the other.
Shirt brightness decides the team
Every shirt measured on the goal window.
V 0, darkshirt brightness (HSV value)V 255, light
dark kits, centre V 52light kits, centre V 248between, needs the vote
Two groups, far apart. The rare values in between are boxes caught mid-duel or half out of frame, which is why each tracked player votes over all of its frames and needs 70% agreement before it gets a team.

Roles (LB, LCB, RDM and so on) come from a Hungarian assignment of detected players to a formation template, frame by frame, smoothed over time. A resolver gives each role to at most one track, the one with the most consistent votes. That uniqueness is enforced, not measured; role accuracy has no hand labels yet.

Twelve seconds at normal speed, one team labelled. A track whose votes have not settled stays unlabelled.

Stage 2: ball detection

Every downstream decision depends on the ball, and at 1080p it is about 10 pixels wide. A rendered ball has its own size, texture and blur, so the ball model is trained on FC footage.

Input resolution. A YOLO11s trained and run at 1280 px. At the common 640 px default the ball shrinks to about 5 pixels and disappears from the feature maps.

Labels. Generated automatically, then audited. A ball-colour mask intersected with a pitch mask gives candidate blobs; blobs are linked across frames into tracks and scored, with a feature that separates the ball from a boot of the same colour; every window is then audited on contact sheets, and about 40% were rejected. The result is 2,530 labels across 45 windows, split by window, not by frame, because consecutive frames are near-duplicates and a frame-level split leaks evaluation frames into training.

Selection is temporal. A look-alike, such as a second ball or a boot, can outscore the true ball for a frame. The picker takes the strongest candidate within a distance budget of the previous position, the budget grows with each missed frame, and a distant candidate takes over only if it persists for three frames.

Twenty seconds of open play at normal speed. The yellow ring is the ball the tracker keeps, its trail fades from green to yellow over the last 25 frames, and grey rings are the other candidates the model returned and continuity rejected.

On 526 hand-audited frames from 9 windows held out of training, none inside the clip shown, with a hit defined as within 20 px of the label: 524 hits (99.6% recall), no wrong object, 0.50 px median error. The most confident detection alone scores 523: continuity removes the one look-alike, and the two misses are frames with no detection.

Stage 3: pitch registration

Every tactical statement is a statement about distances, and distances need a mapping from image pixels to pitch metres: a homography, estimated per frame from landmarks whose pitch coordinates are known.

A YOLOv8s-pose model predicts 32 pitch keypoints (corners, box corners, halfway line, centre circle, penalty arcs). It was trained on a public broadcast dataset (Roboflow, CC BY 4.0) and runs on FC footage as is. RANSAC fits the homography from confident keypoints; keypoints within 2 px of the frame edge are rejected, because off-screen landmarks come back clamped to the border with high confidence.

FC also cuts to a low camera behind the penalty area, which the public dataset never shows. There the model reads the penalty box as the halfway line at full confidence, and the fit lands about 32 m off while passing every self-consistency check. The pipeline catches those frames rather than repairing them: the goal-line hoarding descending into the frame must project onto a goal line, or the frame loses its homography. The repair, a fine-tune on FC frames labelled by hand in a small tool, is measured offline and not yet in the pipeline: four box corners are clicked, and the pitch fitted from them has to land on every painted line, including the ones never clicked.

The labelling tool: a box-camera frame with four clicked box corners and the fitted pitch drawn in magenta over the painted lines, and a side panel listing the next point to click.
The labelling tool. Four box corners clicked by hand; the magenta pitch fitted from them lands on the small box and the arc too, which were never clicked. A wrong click shows up immediately as a line off the paint.

Every frame is graded good, coarse or reject from the fit's own signals (inlier count, a residual under 3 m, freshness) and one external check, run only where the goal-line hoarding is in view. Distance verdicts require good; reconstruction uses any frame that is not rejected. On the match shown, 56.8% of gameplay frames grade good, 11.7% coarse and 31.5% reject. A good grade is self-consistency, not accuracy: a fit can explain its own seven keypoints to 0.5 m and be more than 30 m off the hand-labelled landmarks.

Stage 4: camera-motion compensation

Overlays anchored to the pitch must stay fixed while the camera pans. Each method was scored the same way, by warping the pitch markings 45 frames ahead and measuring the distance to the lines actually observed:

Holding the drawing still
Pitch paint warped 45 frames ahead with each method's camera motion, on 28 frame pairs from one goal build-up. Median distance to the real lines, lower is better.
No motion
39.7 px
ORB features
38.2 px
Median foot shift
25.6 px
Flow, affine, whole frame
20.3 px
Flow, affine, pitch only
10.5 px
Flow, homography, pitch only
1.9 px
Six ways to estimate camera motion, one score. The winner tracks grass corners only, with players, crowd and HUD masked, checks every point forward and backward, and fits a full homography: 1.9 px median, 5.8 px p90.

The tracked player's trajectory is then smoothed offline with a zero-phase local-linear fit, so a trail point stays on the same patch of grass as the camera moves.

Grid comparing the old hologram trail with the rebuilt one following the run on the ground.
The grid that approved the change. Left: the trail carried by the foot-median shift, cut short and dragged toward the player. Right: optical-flow camera, smoothed track, dashes fixed to the ground.

Stage 5: event reconstruction

Passes come from ball ownership. The owner is the nearest player within 1.8 m with a clear margin over the next nearest; a change of owner is a pass candidate, refined to the frame the ball left one player and reached the next. Candidates are rejected if they cross a scene cut, last less than 0.08 s or more than 3 s, or if less than 80% of the flight was observed. A ball that rolls past a teammate belongs to the teammate who controls it: ownership runs where the ball kept 70% of its speed, turned less than 15° and lasted under 0.4 s are discarded.

Pass intercepted
Pass fixes: Pass intercepted
Switch and cut
Pass fixes: Switch and cut
Two passes from the current build. Left: a pass that went to the opponent, drawn to the feet that really took it. Right: the switch and the green cut line to the spot where the defender reaches the passing lane first.

Of 154 reconstructed passes in a full match, the planner annotates 65: passes under pressure, through traffic, line-breaking passes, passes into the final third, turnovers, and passes where a better option was available.

This is the stage that depends on pitch registration. Ownership is decided in metres, and every later annotation starts from the pass list it produces. Ownership compares the ball with players a metre or two away through the same homography, so an error that shifts the whole pitch, like the box camera's, cancels out. An error in local scale does not, and how often it changes an owner has not been measured. Categories that use absolute position, such as a pass into the final third, carry the full registration error. Stages 6 and 7 choose the defender and trace the run in image space.

Stage 6: defensive decisions

When an opponent passes, the question is which defender could have stepped into the passing lane, and whether the user switched to him.

The rule works in image space, where the evidence is, and in hindsight: the lane is the path the ball actually took, which the player could only anticipate. That path is carried back to the kick frame by the median image shift of the players' feet. A defender qualifies if his foot point projects onto the middle 70% of that path, within 3.5 of his own body heights; the nearest qualifying defender is selected. The choice uses no distance in metres. It can only choose among defenders the detector found, so player recall (0.774, against the auto-labeller) bounds it: a missed defender is never named, and the call goes to the next qualifying defender or to none.

An opponent pass with green chevrons on the defender standing in the passing lane and his run traced behind him.
The switch goes to the defender standing in the passing lane, in the centre circle, and his run is traced behind him.

Every instruction is checked. If the controlled player, read from the game's own selection marker, never becomes the named defender inside the instruction's window, the instruction turns red as missed, and a goal conceded within 8 seconds is linked to it. A contain instruction, a line from the defender to the ball carrier modelled on EA's own coaching footage, ends at the next kick, and the next pass gets its own switch instruction.

Stage 7: attacking decisions

In possession, the game already marks the controlled player, so the highlight goes to the teammate who should have received the ball: the open option that was missed, or the receiver of a key pass, from two seconds before the pass until it arrives. His run is traced whenever he moves at least one body height in half a second, curved runs included, like EA's own run indicator.

In possession, green chevrons on the open teammate with his run traced behind him.
In possession the chevrons sit on the open teammate, not on our own ball carrier, and his run is traced behind him.

An overlay that marks everything communicates nothing. A trace appears only on passes the planner annotates, never on a kickoff or a routine pass, and the highlight marks only the player the instruction refers to. The evaluation gate enforces it: a trace on an unannotated pass fails the render.

Development method

Every change starts as a comparison grid: two to four real frames, the current version beside each candidate, drawn from the same data on the same pixels. A grid takes a minute to build; a full render takes fifteen. The change is chosen from the grid, implemented, tested, rendered once, and checked by the gate.

The gate reads the render's sidecar, a record of every overlay drawn on every frame, and checks it against the saved plan and the detections: every trace vertex within 1 px of the camera-compensated flight line, every highlight on a detected player, every caption inside its event's time window. It checks that the render is faithful to the plan and reports detector health; it does not measure whether the plan's decisions are right.

The gate, on the current build
eval_all.py reads the render's sidecar and checks every overlay against the saved plan and the detections.
PASS frame grid PASS one caption text per frame PASS holds on the pass PASS marks on screen PASS switch advice is self-consistent PASS calls are evidence gated FAIL ball coverage on gameplay frames 89.2% of 51851 FAIL longest gameplay gap 283 frames PASS inter-frame jumps over 150px: 0 FAIL static locks while players move: 10 PASS saved-plan overlay timing, player identity and coverage
Part of the gate's output. The three FAILs are one suite, ball health over the whole match rather than the rendered window.

The three FAILs are ball health over the whole match: including set-piece cameras and occluded frames, the tracked ball covers 89.2% of gameplay frames against a 95% target. The check scores the whole match, so every clip rendered from it inherits the FAILs; the checks on the render itself pass. The FAILs put the match in review status, and raising whole-match coverage to 95% is open work. Per-frame accuracy on labelled frames (99.6%) and whole-match coverage (89.2%) measure different things.

Limitations and next steps