by Victoria Hart
July 24, 2026

Live events move fast. Speakers change, audio feeds drop, and attendees with accessibility needs are counting on your captions to keep up in real time. Whether you're running a corporate conference, a university graduation, or a multi-day exhibition, in-room captioning is no longer optional. It's a baseline expectation.
This guide covers everything AV teams and captioners need to know: how to choose between human captioners and automated ASR, how to get your source audio right, how to deliver captions to the right destinations, and how to set it all up on Line 21, a platform built specifically for this workflow, with full engine transparency and a pay-as-you-go pricing model that removes the guesswork.
What Are In-Room Captions and Why Do They Matter?
In-room captions are real-time text displays of spoken content, delivered to attendees at a physical venue. They might appear on a large screen at the front of a room, on an attendee's personal device via a browser or QR code link, or embedded directly into a broadcast feed.
For Deaf and hard-of-hearing attendees, in-room captions are essential. They also benefit non-native speakers, attendees in noisy environments, and anyone who processes information better when they can read along. Accessibility standards including the UK Equality Act 2010 increasingly reinforce the expectation that live events provide accessible communication support.1
Getting it right requires more than switching on a microphone and hoping for the best. It means making deliberate choices across three areas: your captioning method, your audio source, and your caption destination.
Choosing Your Captioning Method: Human Captioners vs. Automated ASR
The first decision is who, or what, produces the captions. There are two primary approaches, and each has a clear place in a professional live event workflow.
Human Captioners: Stenographers and Respeakers
Human captioners fall into two main categories. Stenographers use specialist keyboards to produce captions at high speed, typically achieving accuracy rates above 98% for prepared content.2 Respeakers listen to the speaker and re-voice the audio into speech recognition software, which then generates the captions. This approach can be more cost-effective while still delivering high accuracy when managed by a trained professional.
Human captioners are the gold standard for complex content: panels with multiple speakers, technically dense presentations, or events where errors would cause genuine harm, such as legal proceedings or medical briefings. The trade-off is cost and logistics. You are coordinating with a person, often remotely, and they need the right platform to push captions to your destinations efficiently.
Line 21 is built to make this straightforward. Human captioners, whether stenographers or respeakers, can connect to Line 21 and push live captions directly to HLS streams, browser pages, overlay outputs, RTMP destinations, and CEA 608/708 feeds, without the heavy monthly software outlay that many legacy platforms require. For captioners, it opens up delivery destinations that were previously out of reach.
Automated Captioning (ASR)
Automated Speech Recognition (ASR) captioning uses AI engines to transcribe audio in real time, without a human in the loop. Accuracy varies depending on the engine, the speaker's accent, the audio quality, and the vocabulary involved. For straightforward presentations with clear audio, modern ASR performs well and continues to improve.
For AV teams managing multiple rooms or running events over several days, ASR is a practical default. Line 21 describes it as a "click and forget" service that AV teams can run all weekend without constant oversight. You set it up, connect your audio source, select your engine, and the captions flow.
Line 21 also exposes all the ASR engines it deploys. Rather than locking you into a single provider's model, you can choose which engine to use per language, per session, or based on your own benchmarks. That level of transparency is rare in the captioning platform market, and it matters when you are working across multilingual events or want to compare quality across providers.
Step 1: Getting Your Source Audio Right
Captions are only as good as the audio they come from. Whether you are working with a human captioner or an ASR engine, poor audio quality is the single biggest cause of captioning errors at live events.
Audio Input Options
There are several ways to feed audio into a captioning platform:
- Direct audio feed from the mixing desk: The cleanest and most reliable option. A line-out feed from your audio console sends a processed, mixed signal directly to Line 21, bypassing room acoustics entirely.
- Browser-based microphone input: For smaller or simpler setups, audio can be captured via a browser tab, though this is more susceptible to ambient noise and works best in controlled environments.
- RTMP or HLS audio ingest: If you are already streaming the event, you can route the stream's audio directly into Line 21 for captioning.
Best Practices for a Clean Audio Feed
A few practical rules that make a real difference:
- Use a dedicated, pre-fader output from your desk so that levels stay consistent regardless of what the front-of-house engineer is doing.
- Check microphone gain staging is correct before the session begins. Clipped or underdriven audio degrades ASR accuracy significantly.
- For panels with multiple speakers, use individual lapel or podium microphones rather than a single room mic. Mixed-source audio confuses speech recognition engines.
- Brief human captioners on technical vocabulary, speaker names, and acronyms before the session. Even a short glossary shared in advance can noticeably improve output quality.
- Run a full end-to-end audio test before doors open. Confirm that captions are flowing to all destinations and that latency is within acceptable bounds, typically under three seconds for live events.
Step 2: Choosing Your Caption Destination
Where captions appear matters just as much as how they are generated. Line 21 supports multiple simultaneous caption destinations, so you can serve different audiences and technical requirements from a single session.
On-Screen Display
The most visible caption destination is the venue's display infrastructure: screens at the front of the room, side screens for larger auditoriums, or dedicated caption monitors placed near the audience seating area. Line 21 supports overlay outputs, which allow captions to be rendered directly onto a video feed or display output without requiring a separate screen. This is particularly useful for hybrid events where the in-room and remote experiences need to stay in sync.
Browser-Based Caption Pages
Line 21 generates browser-accessible caption pages that attendees can load on any device, whether a smartphone, tablet, or laptop. No app download is needed, and no account is required. Attendees navigate to a URL and the captions appear in real time. For events where venue screens are limited or sightlines are difficult, browser delivery is a practical and scalable option.
QR Code Access for Attendees
Line 21's browser-based pages can be shared via QR code. Print the code on the event programme, display it on the venue screen at the start of each session, or include it in a delegate app. Attendees scan and read, with no friction and no technical support required. This approach is increasingly popular at accessible-by-design events and fits well with the broader move toward device-agnostic accessibility.
HLS, RTMP, and CEA 608/708 Outputs
For broadcast-grade and hybrid event workflows, Line 21 supports:
- HLS (HTTP Live Streaming): Captions embedded in the HLS stream, accessible to video players and streaming platforms that support the format.
- RTMP: For pushing captioned video to platforms such as YouTube Live, Vimeo, or a custom RTMP endpoint.
- CEA 608/708: The closed caption standard used in broadcast television and many professional video workflows. CEA 608 covers analogue signals; CEA 708 handles digital. If your event output needs to meet broadcast compliance standards, this is the format to use.
These outputs mean that a single Line 21 session can simultaneously serve in-room attendees, remote viewers, and broadcast archives, without duplicating your setup.
Step 3: Setting Up with Line 21
Line 21 is designed to be straightforward from the start. Here is what a typical setup looks like for an AV team or captioner approaching the platform for the first time.
Platform Walkthrough for AV Teams
Once your account is active, you create a session within the platform and configure your audio input. If you are using ASR, you select your preferred engine from Line 21's transparent engine menu. You can see exactly which provider's technology underpins each option. You then configure your output destinations: browser page, overlay, HLS, RTMP, or CEA 608/708. Multiple destinations can run simultaneously from a single session.
The session generates a shareable URL for browser-based access and a QR code you can download straight away. From there, it is a matter of confirming your audio is flowing and monitoring the caption output before you go live.
For AV teams managing multi-room events, separate sessions can be created for each space, with different engines and output configurations per room. This makes Line 21 a practical choice for conference centres and exhibition venues running parallel programming across a full day or weekend.
Platform Walkthrough for Human Captioners
Human captioners connect to Line 21 via their existing tools, whether stenography software or a respeaking setup, and route output into the platform. From there, Line 21 handles delivery to whichever destinations the client has configured. Captioners do not need to manage destination technology themselves. Line 21 takes care of that side of things, so captioners can focus on accuracy rather than technical logistics.
Engine Transparency and Choice
One of Line 21's most distinctive features is its commitment to engine transparency. Most captioning platforms use a single ASR engine and give clients no visibility into which technology is doing the work. Line 21 exposes all the engines it deploys, letting clients choose based on language support, accuracy benchmarks, or prior experience. This is particularly valuable for multilingual events, where different engines can perform better across different languages.
Costs: PAYG Pricing, Transparency, and No Lock-Ins
Captioning platforms have historically run on monthly subscription models that bundle features, lock clients into annual contracts, and charge for capacity regardless of actual usage. For AV teams and captioners who work event by event, this model creates unnecessary cost and commitment.
Line 21 operates on a pay-as-you-go (PAYG) model. You pay for what you use, when you use it, with no minimum monthly spend, no annual contract, and no hidden fees. For a one-day conference, you pay for one day. For a three-day exhibition, you pay for three days. Pricing is visible before you commit.
This matters for event-by-event operators in particular. There is no need to absorb the cost of a software subscription during quiet periods, and no pressure to maximise usage just to justify a fixed outlay. Line 21's model aligns cost directly with activity, which makes budgeting straightforward and predictable.
Frequently Asked Questions
What is the difference between CART captioning and ASR captioning?
CART (Communication Access Realtime Translation) refers to human-produced real-time captions, typically provided by a trained stenographer. ASR (Automatic Speech Recognition) refers to AI-generated captions produced without a human captioner. CART generally offers higher accuracy for complex content; ASR is faster to deploy and more cost-effective for straightforward sessions.
Can in-room captions be delivered in multiple languages simultaneously?
Yes. Line 21 supports multilingual caption delivery, with different language outputs routed to different destinations or displayed simultaneously. Engine selection per language lets you optimise accuracy for each target language.
What audio setup does Line 21 require?
Line 21 accepts audio via browser-based microphone input, RTMP or HLS ingest, and direct audio feeds. For best results, a clean line-out feed from the mixing desk is recommended. Full setup documentation is available in the Line 21 Knowledge Base.
Does Line 21 work for hybrid events?
Yes. Line 21's simultaneous output to HLS, RTMP, browser pages, and overlay means that in-room and remote audiences can receive captions from the same session, with no duplication of effort.
How does PAYG pricing work on Line 21?
You pay based on usage, with no monthly minimums or annual commitments. Pricing is displayed transparently before you begin a session, so there are no unexpected charges.
Is Line 21 suitable for small events as well as large conferences?
Yes. Line 21 is designed to scale. A single-room half-day seminar and a multi-room three-day conference both run through the same platform, with session configuration adapted to the scope of each event.
What caption formats does Line 21 support?
Line 21 supports HLS, RTMP, CEA 608/708, browser-based delivery, and overlay output. Multiple destinations can be active simultaneously within a single session.
Start Delivering In-Room Captions with Line 21
In-room captioning does not have to be complicated or expensive. With the right audio setup, the right captioning method for your content, and a platform that delivers to every destination your event requires, it becomes a reliable and repeatable part of your production workflow.
Line 21 is built for exactly this: professional-grade captioning infrastructure, transparent engine choices, scalable destination support, and pricing that reflects what you actually use. Whether you are an AV team running ASR across a full event weekend or a human captioner looking for a platform that expands your delivery options, Line 21 removes the barriers that have historically made accessible live events harder than they need to be.
Visit line-21.com to explore the platform, or head to the Line 21 Knowledge Base for full setup documentation and guides.
Footnotes
-
UK Equality Act 2010, Section 20: Duty to make reasonable adjustments. legislation.gov.uk ↩
-
National Court Reporters Association (NCRA) sets a minimum accuracy standard of 95% for certified real-time reporters; working accuracy in practice frequently exceeds 98%. ncra.org ↩