Apollo 13's Comm Loops (1970)
When an oxygen tank exploded on Apollo 13, NASA's Mission Control relied on a strict communication structure:
- CAPCOM was the only voice to the crew.
- Specialists (EECOM, GUIDO, etc.) debated on internal loops and routed decisions through the Flight Director.
- Each transmission had a clear role and priority, helping the right information reach the right people during a crisis.
That structure did not change what was said. It changed how information was routed under pressure.
Chat-based language models use a similar structure. A minimal version looks like this:
<SYS> director's notes (policies, style)
<USR> astronaut's question
<AST> ground's reply
These tags are tokens in a single sequence, but they tell the model who is speaking and which instructions belong to each role.
What the Format Buys Us
Clear boundaries improve training. In supervised fine-tuning, loss is often computed only on the assistant span. Role tags make it clear where the model's response begins.
Role tokens provide landmarks. Special tokens such as <SYS> and <USR> give the model consistent boundaries to attend to, including in long prompts.
Consistent templates reduce ambiguity. Using the same template for training and production helps the model learn the structure it will encounter at inference time.
Diagram: One Stream, Labeled Spans
Alt text: A left-to-right token stream with <SYS>, <USR>, <AST> blocks and an arrow showing generation order.
How Formatting Helps Attention
Consider a simplified example with the same content encoded in two ways. The diagrams illustrate how explicit roles can give attention a clearer structure:
With Roles: Attention Peaks at the System Boundary and User Span
Without Roles: Attention Diffuses Across Mid-Sentence Tokens
Why This Happens
Boundary tokens are distinctive. Repeated role markers can develop representations that make them easier for the model to identify than ordinary prose.
The model has less to infer. A <SYS> marker explicitly identifies the span containing system instructions.
Markers remain useful over long contexts. They can help the model locate important spans, although formatting alone does not solve the "lost in the middle" problem.
Training-Time View: Assistant-Only Loss
In a common SFT setup, labels outside the assistant span are masked, so system, user, and role tokens have a loss of zero. Only the assistant's response contributes to the gradient:
Only the assistant's response ("Bleu.") contributes to the training loss
Chat formatting is more than a presentation choice. Like NASA's communication protocols, it gives a stream of information an explicit structure.
For a language model working through a long prompt, that structure can make instructions and conversational turns easier to distinguish.
References
- Patterson, E. S., Watts-Perotti, J., & Woods, D. D. (1999). Voice Loops as Coordination Aids in Space Shuttle Mission Control. Computer Supported Cooperative Work.
- NASA. (1970). Apollo 13 Flight Journal - Day 3: The Problem.
- Hugging Face. Chat Templates. Transformers Documentation.
- Hugging Face. Chat Templates. Hugging Face NLP Course.
- Hugging Face. SFT Trainer. TRL Documentation.
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv preprint.