Human speech
Live events stop at language.
PolyCast
An AZCY platform.
Broadcast one live event in every language your audience speaks, at the same time — and see how each language stream is doing while it runs.
A congregation of five hundred people, four languages in the room, one interpreter serving one of them. Everyone else can hear the service but cannot follow it. The technology to fix this existed — it was just built for conference centres in Geneva and priced accordingly.
The shape of it
What moves, and what happens to it.
One live source
A microphone, a camera, or the production setup you already run. The room does not change what it does — PolyCast takes the feed you already produce.
Listen, translate, speak
The system listens to the speaker, translates into each language, and voices the result — live, moments behind the speaker, not a recording processed afterwards.
Every language, at once
One stream per language, sent to YouTube, Facebook, or a player on your own site. Each stream shows how far behind the room it is.
Capabilities
The parts that carry the argument.
One source, every language, at once
Add a language and it becomes its own live stream. The speaker does nothing differently.
See the delay on every stream
Every language stream shows how far behind the speaker it is, live. If one drifts, you see which one, and by how much.
Interpreter-assist mode
A human interpreter corrects the machine's translation live, rather than translating from scratch — put your interpreter on the language the machine handles worst.
African language packs
Yorùbá, Igbo, Hausa, and Nigerian Pidgin among them. This is the part the big international tools do not cover.
Per-language archives
Every language is recorded separately, so the Yorùbá service exists afterwards as its own recording, not buried in a mix.
How it fits
Into what you already run.
Nothing here replaces a system that works. These are the questions that decide whether a platform can be deployed at all, so they are answered before the specification.
Integrations
- OBS and software encoders
- Hardware encoders over SRT/RTMP
- YouTube Live and Facebook Live
- Custom RTMP endpoints
- Embedded player for an existing site
Deployment
- Managed cloud
- Regional deployment to shorten the path from the venue
- On-premise for fixed broadcast installations
Security and custody
- Per-session access control
- Private channels for internal or paid events
- Archives are retained on the operator's schedule, not ours
Specification
Design facts and stated targets. Nothing we have not measured.
Technical specification
Ingest
- Protocols
- SRT · RTMP
- Sources
- Microphone · Camera · OBS · Hardware encoder
- Source model
- One source, N language channels
Pipeline
- Latency budget
- Design target: under 1s end to end
- Mode
- Machine pass · Interpreter-assist
- Output
- Text · Synthesised voice
- Language isolation
- Per-channel, independent
Distribution
- Destinations
- YouTube · Facebook · Custom RTMP · In-house player
- Channels
- One per target language
- Recording
- Per-language archive
- Monitoring
- Per-channel latency and health
The hard problem
Translation quality and latency are the same dial, and a live room only forgives one of them.
In plain terms: the hard trade is between translation quality and delay — every moment spent getting a better translation is a moment your audience falls further behind the room. Every accuracy gain in this pipeline is bought with time. Wait for more audio and the recognition is better, the translation has more context, and the sentence is more likely to be right — and the listener is now far enough behind the speaker that they are watching a different event. Under about a second, the room stays together. Past a couple of seconds, the audience in the translated channel stops experiencing a live event and starts experiencing a delayed one, and they can hear the original in the room, which makes the gap worse rather than better.
So the pipeline is built against a latency budget, and every stage is allocated a slice of it. Recognition segments on natural boundaries where it can and forces a boundary where it must, because waiting indefinitely for a speaker to finish a sentence is how a live channel falls behind and never recovers. Translation is per language and independent — a slow language is a slow channel, never a slow broadcast. That isolation is the single most important structural decision in the system: without it, adding a fifth language degrades the other four, and the failure looks like the platform is bad rather than like one language pack is.
Then the languages themselves. The commercial pipelines are excellent at French and German and thin on Yorùbá, Igbo, and Hausa — which is precisely inverted from what a Nigerian broadcast needs. That gap is why interpreter-assist exists. A human interpreter working over a machine pass is doing correction rather than origination: a far lower cognitive load than booth interpreting, sustainable across a long service, and it puts a competent human exactly where the machine is weakest. The design admits what the models cannot do yet, rather than pretending.
The unglamorous half is the connection. Venues have the network they have. SRT exists because RTMP over a bad link degrades in ways an audience notices immediately, and SRT recovers packet loss at the cost of a little more latency — which then has to come out of a budget that was already tight. That trade is made per venue, not once in an architecture document.
The practices behind it
Where it runs
Questions
The ones we actually get asked.
Do we need to change how we run our production?
No. If you already run a production desk or streaming equipment, PolyCast takes the output you already produce (over SRT or RTMP). The people running the desk keep doing what they do.
How many languages can one session carry?
Each language runs as its own stream, so the practical limit is your venue's internet and the cost you choose to carry — not a number fixed in the platform.
Is the machine translation good enough on its own?
It depends entirely on the language pair, and we will tell you which of yours are strong and which are not. Where the machine is weak, interpreter-assist puts a person over its output to correct it live. We would rather tell you that honestly than have you find out during a service.
Can we keep the recordings?
Yes, per language. Each channel is archived as its own recording rather than as a mix, so the Yorùbá stream exists afterwards as the Yorùbá stream.
What happens if a destination drops mid-session?
That channel reports unhealthy and the others keep running. The operator sees which destination failed while it is happening, rather than discovering it in the comments afterwards.
We build systems like this for other organizations.
PolyCast exists because we met the same operational failure enough times to engineer it properly. Yours is probably a different failure. That is the conversation we want, and you will have it with an engineer.