Latency Budget: Where the Milliseconds Go
Latency is not one number a vendor gives you. It is a budget you allocate across five stages, and the stage you forgot is usually the one that breaks it.
- Glass to glass: the five stages
- Stage 1: capture and encoding
- Stage 2: transport and buffering
- Stage 3: the network itself
- Stage 4: decoding
- Stage 5: the display
- Allocating the budget
- Keeping several screens on the same moment
- How to measure it
- FAQ
Glass to glass: the five stages
Glass-to-glass latency is the delay between something happening in front of the camera and appearing on the screen. It decomposes into five stages, and every one of them is separately adjustable:
- Capture and encoding — compression and the buffer it requires.
- Transport and buffering — the protocol's framing, segmentation and, where it retransmits, its recovery buffer.
- The network — switching, routing, queuing, and distance.
- Decoding — the receiver's buffer and decode time.
- The display — processing inside the screen itself.
The reason to break it down is that the stages trade against each other. Reducing the recovery buffer makes the stream more fragile on a lossy path; raising it makes an interactive application unusable. There is no setting that minimises latency and maximises robustness at the same time.
Stage 1: capture and encoding
Compression costs time, because the encoder has to hold frames in order to reference them. Structure matters: a longer group of pictures compresses better at the same quality but adds delay, and bidirectional prediction adds more again. A low-latency profile shortens or removes that dependency at the cost of bitrate.
Hardware and software encoding differ here, and hardware is generally the more predictable choice for live work — the encode is dedicated silicon with a fixed pipeline rather than a share of a general-purpose CPU. The comparison is set out in Hardware vs Software Encoding.
Where the encoder gives you a latency-related setting, take it, but verify the result: a setting described as low latency is a starting point, not a guarantee about the whole chain.
Stage 2: transport and buffering
This is where the protocol choice lands, and the differences are structural rather than incremental.
- UDP and RTP on a LAN have minimal framing and no recovery buffer, which is why they are the low-latency default for local distribution.
- RTSP adds session control over the top and is common for local monitoring; details in What Is RTMP? for the push model and the streaming setup guide for practical configuration.
- RTMP and RTMPS are what platforms accept for ingest, and carry more buffering than raw UDP. They are the right answer when the destination is a public platform, regardless of latency.
- HLS and HTTP-FLV trade latency for compatibility, which is why they suit broad browser and mobile reach rather than live interaction.
- SRT is the interesting one because its recovery buffer is configurable: you choose how much loss it absorbs, and you pay for that in delay. On a clean path, configure it tight; on a lossy public path, loosen it and accept the delay. Background in What Is SRT?.
- WebRTC is built for interactive use. The ZY-EH1401 supports standard WebRTC with transmission latency quoted at approximately 300 ms, which is the figure to design against for two-way or interactive work.
The rule: pick the protocol by the destination and the path, then spend the latency budget where it actually buys something. Choosing RTMP because a platform requires it is not a latency decision; choosing an unnecessarily large SRT buffer on a clean LAN is.
Stage 3: the network itself
Propagation delay over distance is usually small compared with the delays introduced by equipment. What costs time is queuing: a congested link, a buffer that is too large, or a wireless hop that retransmits at layer 2.
- Keep it switched, not routed, where the topology allows. Every router hop adds processing.
- Avoid congestion by sizing the link properly — a saturated link queues, and queuing is latency. See How Much Bandwidth Does an IP Video Link Need?.
- Do not put interactive video through a wireless hop unless the application can tolerate it. Wireless retransmission is invisible in throughput tests and obvious in latency.
- Prefer multicast for one-to-many, so the source sends once rather than queueing one copy per receiver.
Stage 4: decoding
The decoder's own buffer is often the largest single contributor in a chain nobody examined. It exists to absorb jitter, and it is frequently left at a conservative default. Where the application is interactive, reduce it and confirm the picture holds on the real network, not on a bench.
Decoder choice matters here as much as encoder choice. The ZY-DH901 is specified at under 200 ms, and matching decoder models across a site is worth more than optimising one of them.
Stage 5: the display
Modern displays add their own processing — scaling, deinterlacing, motion smoothing — and it is not small. A consumer television in its default picture mode can add more delay than the entire IP chain. Where latency matters, put the display in a low-latency or game mode, or specify a monitor intended for the purpose.
This stage is routinely overlooked because nobody considers the screen part of the video system. It is, and it is often the cheapest latency reduction available.
Allocating the budget
| Application | What dominates | Where to spend the budget |
|---|---|---|
| Live production, camera-to-wall | Encode plus decode buffers plus display | Shorten decoder buffer, low-latency display mode, keep it on a LAN |
| Interactive or two-way | Transport buffer | WebRTC where available; otherwise tighten SRT and accept less loss protection |
| Platform streaming | Platform ingest and its own delay | Accept it — the protocol is set by the destination |
| Multi-screen display | Consistency between paths, not absolute value | Same source, same protocol, same decoder model, multicast |
| Recording and archive | Nothing | Latency is irrelevant; optimise quality and storage instead |
Write the budget down before commissioning. Deciding during an event that latency is unacceptable is a conversation nobody wants to have.
Keeping several screens on the same moment
Synchronised screens are a consistency problem, not a minimum-latency problem. Screens drift because their paths differ, not because any one path is slow. Four things hold them together:
- One source, one stream. Feeding screens from the same multicast group means they receive identical packets at the same instant.
- Same decoder model throughout. Different models have different buffer defaults, and that difference shows as visible drift between adjacent screens.
- Same display model and same picture mode. Differing display processing is the most common cause of adjacent screens being visibly out of step.
- Multicast rather than parallel unicast. Parallel unicast streams start at different moments and drift independently.
Then verify visually: put a running clock or a fast-moving test pattern on the source and photograph the screens together. Tolerance is whatever the application requires, and it should be stated before commissioning rather than argued afterwards.
How to measure it
The reliable field method needs no instruments. Point a camera at a running clock, send that picture through the chain, and photograph the clock and the display together in one frame. The difference is your glass-to-glass figure, and it includes every stage.
Do it at the start and record it, then repeat after any change to protocol, buffer or display. A single recorded number settles most later disputes about whether something got slower.
Where to start
ZY-EH1401 4K Encoder
Standard WebRTC support with transmission latency quoted at approximately 300 ms for interactive work.
InteractiveZY-DH901 Multi-Interface Decoder
Specified at under 200 ms — the receiving end of a latency budget.
Low-latency decodeZY-EH1304 4-Channel 4K Encoder
Four channels on one multicast source, so screens receive identical packets at the same instant.
Synchronised screensZY-EH1308 8-Channel 4K Encoder
Eight channels with two network ports for separating distribution from management traffic.
Higher densityFAQ
What latency should I expect from an IP video chain?
It depends on which stages you control. Capture and encoding, transport buffering, the network, decoding and the display each contribute, and the decoder buffer and display processing are often the largest. Where every stage is optimised for it — a LAN, minimal buffering, a low-latency display mode — the result can be low enough for interactive work. Where the destination is a public platform, its ingest and its own delay dominate and no encoder setting will change that.
Which protocol has the lowest latency?
On a local network, UDP or RTP, because there is minimal framing and no recovery buffer. For interactive use across a network, WebRTC is designed for it — the ZY-EH1401 quotes approximately 300 ms for WebRTC transmission. SRT sits in between and is the useful choice because its recovery buffer is configurable: tighten it on a clean path, loosen it on a lossy one. HLS trades latency for reach and is the wrong choice for interaction.
Why are my screens showing different moments?
Because their paths differ, not because any path is slow. Different decoder models have different buffer defaults, different displays add different processing delay, and parallel unicast streams start at different times. Feed them from one multicast group, use the same decoder model and display model throughout, put displays in the same picture mode, and verify with a fast-moving pattern photographed across the screens together.
Can I reduce latency by lowering the bitrate?
Not directly. Bitrate affects quality and bandwidth, not the delay added by compression structure or buffering. To reduce encoding latency you change the encode structure or the low-latency setting; to reduce transport latency you change protocol or buffer size. Lowering bitrate may help indirectly where congestion was causing queuing, but that is a network fix rather than an encoding one.
Does my display really matter?
Yes, and it is usually the cheapest fix available. Consumer televisions apply scaling, deinterlacing and motion processing in their default picture mode, and that processing can exceed the delay of the whole IP chain. Put the display in a low-latency or game mode, or specify a monitor intended for the purpose, before spending money anywhere else in the chain.
Related reading
Need a latency figure before you commit? Send us the application and topology
30-day free trial | 24/7 technical support | 3-year warranty
Get a recommendation 400-056-8185