ECCV 2026

The Telephone Game: Evaluating Semantic Drift in Unified Models

How much meaning survives when a model turns text into images, and back again?

Institute of Artificial Intelligence, University of Central Florida, USA

* Corresponding author · † Equal second authors · ‡ Equal third authors

01 / The evaluation gap

Unified models need cyclic evaluation.

Unified models can both describe images and generate them. Standard benchmarks test these abilities separately; our telephone game tests what happens when a model repeatedly uses them together.

A unified model alternates text and image generation from a suitcase and a banana. The suitcase disappears after generation five, while the banana count grows to fourteen.
The chain begins with “a suitcase left of a banana.” As the model alternates between generating images and describing them, the suitcase disappears and one banana becomes fourteen—revealing how meaning drifts across generations.

What a single pass misses

  • Generation benchmarks assess realism or prompt compliance in one image.
  • Understanding benchmarks assess how well a model interprets an image.
  • Cyclic evaluation tests whether the same model preserves meaning when these capabilities are composed.
BAGEL identifies a winning chess position for white, but generates a generic board when asked to depict white winning.

The gap in one example

Understanding a concept does not guarantee faithful generation.

BAGEL correctly identifies a winning chess position, yet generates a generic board when asked to depict the same concept. This mismatch is cross-inconsistency.

Watch & explore

The presentation and poster.

Preview of the Telephone Game research poster with the evaluation protocol, semantic drift examples, and model results.

Conference poster

A visual overview of the work.

View PDF

02 / The protocol

Follow the meaning through the whole chain.

The Semantic Drift Protocol (SDP) starts with a caption or an image, then alternates text-to-image (T2I) and image-to-text (I2T). Each output becomes the next input; fidelity is always measured against the original.

Text-first chain

TextImageText

Does the original description survive?

Image-first chain

ImageTextImage

Does the original scene survive?

Text-first and image-first chains, with comparisons back to the original input. A red truck changes identity and position, while an image of a group gains people and new objects.
Text-first and image-first chains expose changes that a plausible individual output can hide. Comparisons to the original input track retention within and across modalities.

Two complementary views of semantic retention

Embedding similarity captures the scene’s overall meaning; object-level evaluation checks its specific constraints.

MCD ↑ Higher is better

Mean Cumulative Drift

Average semantic similarity across generations and four modality mappings.

  • MPNet for text ↔ text
  • DINO for image ↔ image
  • CLIP for cross-modal similarity

Despite the name, a higher MCD score means stronger retention and less drift.

MGG ↑ Higher is better

Multi-Generation GenEval

Object-level prompt compliance across repeated generations, using rewritten GenEval prompts and OWLv2 detections.

ObjectsCountsColorsPositionsAttribute binding

Measures whether a model preserves structured scene details over time.

03 / Qualitative examples

Small changes become a different scene.

Repeated generation can change an object’s identity, position, count, color, or style—and introduce content that was never in the prompt. These six examples show semantic drift unfolding across the chain.

Six rows of generated images show position errors, a bat turning into a spoon, a photograph turning into a cartoon, lost clock counts, a hallucinated city, and a bus changing color.
Drift affects entities, attributes, counts, relations, and style—even when individual outputs look plausible.
01

Spatial relations

The clock’s position is lost.

02

Object identity

A baseball bat becomes a spoon.

03

Visual style

A photograph becomes a cartoon.

04

Object counts

Four clocks become an unspecified set.

05

Hallucinated content

A city appears around an empty road.

06

Color attributes

A brown bus turns yellow.

04 / Comparison across models

Every model drifts. Some preserve much more.

BAGEL leads both retention measures across the seven systems. LLaVA+SDXL illustrates why both metrics matter: preserving individual objects does not guarantee that the scene’s overall meaning survives.

Seven systems · camera-ready results
ModelMCDavg ↑MGG ↑
BAGEL Best0.441478.2
BLIP3-o0.398462.2
Show-o0.369157.8
Janus-7B0.377852.5
LLaVA+SDXL0.329832.3
Janus-1.3B0.347027.0
VILA-U0.282919.7

Higher is better for both metrics. Rows follow the paper’s MGG ordering; MCD rankings differ.

GenEval accuracy across generations. BAGEL remains comparatively stable, while weaker models decline sharply.
Similar first-generation compliance can conceal very different long-term behavior.

05 / Across generations

A strong first pass can hide a fragile chain.

Similar initial scores diverge as models repeatedly use their own outputs. BAGEL remains comparatively stable, while weaker models lose their connection to the original scene.

MPNet similarity between initial and later captions in text-first chains. BAGEL preserves the most meaning; VILA-U declines steeply.
Text → text · MPNet semantic similarity
DINO similarity between initial and generated images in image-first chains. Visual fidelity declines as generations progress.
Image → image · DINO semantic similarity
Explore the cross-modal results
CLIP similarity between original captions and generated images across text-first chains.
Text → image · CLIP semantic similarity
CLIP similarity between original images and later captions across image-first chains.
Image → text · CLIP semantic similarity

06 / What fails first

Objects can survive after their relationships are lost.

Simple object presence and color persist longer. Counts, spatial relations, and attribute bindings tend to fail earlier.

Generation when each property is first lost · higher means longer retention
ModelSingle objectTwo objectsCountPositionColorBinding
BAGEL201510111917
VILA-U10222102

Two representative systems from the poster’s category-persistence comparison.

Once a core detail is lost, it rarely returns.

A critical mistake changes the input for every later step. Subsequent generations compound the error, making semantic collapse difficult to recover from.

07 / Human evaluation

The drift is also visible to people.

Six annotators assessed understanding and generation on 200 examples, with two independent annotations per example. Model identities were hidden and output order randomized. MCD and MGG strongly agree with their rankings.

MCD vs. understanding rankr = −0.833

Strong agreement with human judgments.

MGG vs. generation rankr = −0.821

Closer agreement than single-pass GenEval.

GenEval vs. generation rankr = −0.753

The single-pass comparison.

Pearson correlations across the seven systems. Negative values reflect that lower human rank is better, while higher metric scores are better.

Single-pass GenEval versus human generation ranking, with Pearson r of minus 0.753.
Single-pass GenEval vs. human generation rank
Multi-Generation GenEval versus human generation ranking, with Pearson r of minus 0.821.
Multi-Generation GenEval vs. human generation rank

Understanding succeeds more often than generation.

The human study reveals the same mismatch that motivates cyclic evaluation: a model can understand an input well but fail to generate it faithfully.

Human fidelity ratings show stronger understanding than generation across most systems. A detailed BAGEL heatmap shows fewer mismatches.
Good understanding paired with poor generation is a common failure. BAGEL shows the strongest cross-consistency among these systems.
Cross-consistency failures in the human study · percentage of examples
ModelGood understanding,
poor generation
Poor understanding,
good generation
BLIP3-o22.7%2.2%
Janus-7B13.0%0.5%
Show-o6.5%4.3%
Checks beyond the main benchmark
  • Representation checks: alternative embedding backbones preserve the main trends.
  • Prompt sensitivity: different recaptioning prompts retain the relative conclusions for BAGEL, Show-o, and VILA-U.
  • Detector audit: OWLv2 achieves 81.0% agreement with human judgments; agreement remains similar at early and late generations.
  • Additional models: Emu3 and proprietary systems are evaluated on smaller diagnostic subsets, separately from the full benchmark ranking.

Use this work

Citation

Accepted at ECCV 2026. The BibTeX below references the arXiv preprint.

@misc{mollah2025telephonegameevaluatingsemantic,
  title={The Telephone Game: Evaluating Semantic Drift in Unified Models},
  author={Sabbir Mollah and Rohit Gupta and Sirnam Swetha and
          Qingyang Liu and Ahnaf Munir and Mubarak Shah},
  year={2025},
  eprint={2509.04438},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2509.04438}
}