How FUT HyperVision works
Ultimate Team is the most competitive mode in EA SPORTS FC. Matches are fast, opponents are good, and games are decided by details: a switch made half a second late, a passing lane left open, a teammate nobody saw. I have played it for years, and I built FUT Evolution, a companion site and app for its players, so I spend a lot of time around people trying to get better at it.
Under almost every match-analysis video on YouTube, the same question comes back: could AI watch a match and break it down the way a coach would, from where the players stand and where the ball goes, and turn that into advice for the next game?
What a coach does on a paused frame is not magic. It is perception, geometry and a handful of football rules, and each of those can be built and measured. I took the problem apart that way, one piece at a time, and named the system that came out of it FUT HyperVision. The name is my own, not an EA feature. This post shows how each piece is built and what it is measured against. Whether the coaching calls agree with an expert is the next measurement, not a result yet.
What it does
FUT HyperVision takes a recorded FC match and returns the same video with tactical annotations rendered into it: which defender should have been selected to cut a passing lane, where he could reach it, which teammate was open, and whether the player followed the instruction. The annotations follow the visual language EA uses in the game, so they read as part of the broadcast rather than as a debug overlay.
Under the hood it is a perception and reasoning pipeline over 1080p, 60 fps video. It detects and tracks the players, the officials and a ball about 10 pixels wide, registers each frame to pitch coordinates and grades the fit, compensates for camera motion, reconstructs passes and ownership, and runs a rule-based planner over the whole match before a single pixel is drawn.
Three constraints shape every stage. The footage is rendered, not filmed, so models trained on real football transfer poorly. There is no ground truth for most layers, so every evaluation has to be built. And the output is advice, so a confident wrong answer is worse than no answer.
| Layer | Result | Measured against |
|---|---|---|
| Ball detection | 99.6% recall, 0 wrong objects, 0.50 px median | 526 hand-audited frames, 9 held-out windows |
| Player detection | mAP50 0.872, precision 0.934, recall 0.774 | 200 held-out auto-labelled images |
| Camera-motion compensation | 1.9 px median at 45 frames | observed pitch markings, 28 frame pairs |
| Pitch registration | Graded per frame: 56.8% good, 11.7% coarse, 31.5% reject | Self-consistency only, see limitations |
The pipeline
Stage 1: player detection
A YOLO26s detector with three classes (player, goalkeeper, referee), run at 960 px and tracked with ByteTrack. The tracker is reset on hard cuts, so a replay does not inherit the identities of the live play it interrupted. Officials and goalkeepers are detected explicitly so the planner can reason about outfield players only.
On a held-out split of 200 images the model reaches mAP50 0.872 overall (0.955 players, 0.927 referees, 0.735 goalkeepers), with precision 0.934 and recall 0.774. That split comes from our own auto-labeller, so these figures measure agreement with it, not accuracy against human labels.
Officials have to be found, not just ignored: a referee standing in a passing lane looks like one more player, and the switch rule must never name him. The referee class covers the assistant referees on the touchline too. Its training labels come from an auto-labeller that runs a generic person detector, assigns the class from kit-colour rules, and keeps a person only when the band under its feet is at least 30% pitch-coloured pixels, which removes spectators behind the advertising boards.
Teams are split by brightness, not hue: a box around a running player is mostly grass, and hue measures the grass. The classifier samples the upper chest (22 to 48% of the box height), masks out grass pixels and keeps the median HSV value. The values separate into two clusters with a one-dimensional 2-means. The user's team is the cluster under EA's own marker for the controlled player. For the match shown, the match configuration sets the team instead, along with the attacking direction of each half. Each track votes over all its frames and needs 70% agreement to get a team.
Roles (LB, LCB, RDM and so on) come from a Hungarian assignment of detected players to a formation template, frame by frame, smoothed over time. A resolver gives each role to at most one track, the one with the most consistent votes. That uniqueness is enforced, not measured; role accuracy has no hand labels yet.
Stage 2: ball detection
Every downstream decision depends on the ball, and at 1080p it is about 10 pixels wide. A rendered ball has its own size, texture and blur, so the ball model is trained on FC footage.
Input resolution. A YOLO11s trained and run at 1280 px. At the common 640 px default the ball shrinks to about 5 pixels and disappears from the feature maps.
Labels. Generated automatically, then audited. A ball-colour mask intersected with a pitch mask gives candidate blobs; blobs are linked across frames into tracks and scored, with a feature that separates the ball from a boot of the same colour; every window is then audited on contact sheets, and about 40% were rejected. The result is 2,530 labels across 45 windows, split by window, not by frame, because consecutive frames are near-duplicates and a frame-level split leaks evaluation frames into training.
Selection is temporal. A look-alike, such as a second ball or a boot, can outscore the true ball for a frame. The picker takes the strongest candidate within a distance budget of the previous position, the budget grows with each missed frame, and a distant candidate takes over only if it persists for three frames.
On 526 hand-audited frames from 9 windows held out of training, none inside the clip shown, with a hit defined as within 20 px of the label: 524 hits (99.6% recall), no wrong object, 0.50 px median error. The most confident detection alone scores 523: continuity removes the one look-alike, and the two misses are frames with no detection.
Stage 3: pitch registration
Every tactical statement is a statement about distances, and distances need a mapping from image pixels to pitch metres: a homography, estimated per frame from landmarks whose pitch coordinates are known.
A YOLOv8s-pose model predicts 32 pitch keypoints (corners, box corners, halfway line, centre circle, penalty arcs). It was trained on a public broadcast dataset (Roboflow, CC BY 4.0) and runs on FC footage as is. RANSAC fits the homography from confident keypoints; keypoints within 2 px of the frame edge are rejected, because off-screen landmarks come back clamped to the border with high confidence.
FC also cuts to a low camera behind the penalty area, which the public dataset never shows. There the model reads the penalty box as the halfway line at full confidence, and the fit lands about 32 m off while passing every self-consistency check. The pipeline catches those frames rather than repairing them: the goal-line hoarding descending into the frame must project onto a goal line, or the frame loses its homography. The repair, a fine-tune on FC frames labelled by hand in a small tool, is measured offline and not yet in the pipeline: four box corners are clicked, and the pitch fitted from them has to land on every painted line, including the ones never clicked.
Every frame is graded good, coarse or reject from the fit's own signals (inlier count, a residual under 3 m, freshness) and one external check, run only where the goal-line hoarding is in view. Distance verdicts require good; reconstruction uses any frame that is not rejected. On the match shown, 56.8% of gameplay frames grade good, 11.7% coarse and 31.5% reject. A good grade is self-consistency, not accuracy: a fit can explain its own seven keypoints to 0.5 m and be more than 30 m off the hand-labelled landmarks.
Stage 4: camera-motion compensation
Overlays anchored to the pitch must stay fixed while the camera pans. Each method was scored the same way, by warping the pitch markings 45 frames ahead and measuring the distance to the lines actually observed:
The tracked player's trajectory is then smoothed offline with a zero-phase local-linear fit, so a trail point stays on the same patch of grass as the camera moves.
Stage 5: event reconstruction
Passes come from ball ownership. The owner is the nearest player within 1.8 m with a clear margin over the next nearest; a change of owner is a pass candidate, refined to the frame the ball left one player and reached the next. Candidates are rejected if they cross a scene cut, last less than 0.08 s or more than 3 s, or if less than 80% of the flight was observed. A ball that rolls past a teammate belongs to the teammate who controls it: ownership runs where the ball kept 70% of its speed, turned less than 15° and lasted under 0.4 s are discarded.


Of 154 reconstructed passes in a full match, the planner annotates 65: passes under pressure, through traffic, line-breaking passes, passes into the final third, turnovers, and passes where a better option was available.
This is the stage that depends on pitch registration. Ownership is decided in metres, and every later annotation starts from the pass list it produces. Ownership compares the ball with players a metre or two away through the same homography, so an error that shifts the whole pitch, like the box camera's, cancels out. An error in local scale does not, and how often it changes an owner has not been measured. Categories that use absolute position, such as a pass into the final third, carry the full registration error. Stages 6 and 7 choose the defender and trace the run in image space.
Stage 6: defensive decisions
When an opponent passes, the question is which defender could have stepped into the passing lane, and whether the user switched to him.
The rule works in image space, where the evidence is, and in hindsight: the lane is the path the ball actually took, which the player could only anticipate. That path is carried back to the kick frame by the median image shift of the players' feet. A defender qualifies if his foot point projects onto the middle 70% of that path, within 3.5 of his own body heights; the nearest qualifying defender is selected. The choice uses no distance in metres. It can only choose among defenders the detector found, so player recall (0.774, against the auto-labeller) bounds it: a missed defender is never named, and the call goes to the next qualifying defender or to none.
Every instruction is checked. If the controlled player, read from the game's own selection marker, never becomes the named defender inside the instruction's window, the instruction turns red as missed, and a goal conceded within 8 seconds is linked to it. A contain instruction, a line from the defender to the ball carrier modelled on EA's own coaching footage, ends at the next kick, and the next pass gets its own switch instruction.
Stage 7: attacking decisions
In possession, the game already marks the controlled player, so the highlight goes to the teammate who should have received the ball: the open option that was missed, or the receiver of a key pass, from two seconds before the pass until it arrives. His run is traced whenever he moves at least one body height in half a second, curved runs included, like EA's own run indicator.
An overlay that marks everything communicates nothing. A trace appears only on passes the planner annotates, never on a kickoff or a routine pass, and the highlight marks only the player the instruction refers to. The evaluation gate enforces it: a trace on an unannotated pass fails the render.
Development method
Every change starts as a comparison grid: two to four real frames, the current version beside each candidate, drawn from the same data on the same pixels. A grid takes a minute to build; a full render takes fifteen. The change is chosen from the grid, implemented, tested, rendered once, and checked by the gate.
The gate reads the render's sidecar, a record of every overlay drawn on every frame, and checks it against the saved plan and the detections: every trace vertex within 1 px of the camera-compensated flight line, every highlight on a detected player, every caption inside its event's time window. It checks that the render is faithful to the plan and reports detector health; it does not measure whether the plan's decisions are right.
The three FAILs are ball health over the whole match: including set-piece cameras and occluded frames, the tracked ball covers 89.2% of gameplay frames against a 95% target. The check scores the whole match, so every clip rendered from it inherits the FAILs; the checks on the render itself pass. The FAILs put the match in review status, and raising whole-match coverage to 95% is open work. Per-frame accuracy on labelled frames (99.6%) and whole-match coverage (89.2%) measure different things.
Limitations and next steps
- New game editions. The detectors were trained mainly on FC 26 footage. A new ball texture, new kits and a new broadcast camera are distribution shifts. The ball detector fails by abstaining up to an 80° hue shift; the pitch model can fail confidently, as it does on the box camera. Each new edition means labelling its frames with the same pipelines and retraining.
- Ball evaluation. The 526 frames are the validation split, which also picked the checkpoint and informed the picker. Labels were seeded by a colour mask, so the hardest windows are under-represented and no frame without a ball is included: "no wrong object" says nothing about false alarms on empty frames. An untouched test set with ball-absent frames is next.
- Pitch registration. Broadcast registration has no reliable accuracy figure yet. The good grade is almost entirely self-consistency: of 29,463 gameplay frames graded good on the match shown, 18 had the hoarding in view for the external check. The only hand-labelled broadcast frame graded good is 3 to 11 m off at its six labelled landmarks. The next step is more hand-labelled broadcast frames, FC-specific keypoint training with the box-camera recipe, and moving the switch lane from the foot-median shift to the flow camera model.
- Events. Ownership and passes are computed in pitch metres on any frame that is not rejected, with the ball projected to the ground even when it is in the air. There is no hand-labelled pass list yet, so pass precision and recall are unmeasured.
- Decision quality. The switch, contain and interception rules are heuristics with measured behaviour, not a model trained on labelled tactical decisions. The next step is labelling what an expert would have done and scoring the rules against it. The followed-or-missed check reads only the cursor: a defender already selected counts as followed, and the game's auto-switch setting is unknown. Team side, attacking direction per half and goal times come from match configuration.