Back to work

NeurIPS 2026

ReSCUE

Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

Live captions for sign language that keep up with a signer for minutes at a time, instead of one pre-cut sentence at a time.

Authors
Sihan Ren*, Gaozheng Li*, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye Wu. *Equal contribution.
Institutions
ShanghaiTech University, University of Derby and DGene. Work done during an internship at DGene.
My part
Wrote the proposal for a real-time system that translates a continuous video stream.
Links
PaperCodeProject page
ReSCUE running on a continuous signing video. The caption for the current sentence keeps changing as more signs arrive, and is locked once the sentence ends. The signer's video was processed with AI to protect privacy, so some movements may look slightly distorted.

The problem

Real-time sign language translation could let a deaf signer and a hearing listener talk without waiting for an interpreter. Most translation systems, though, work offline and one sentence at a time, and they expect someone to have already cut the video into sentences.

Real signing does not arrive that way. It runs on for minutes, with pauses that are not sentence breaks. A live system has to decide what to show before a sentence is over, revise it without the caption flickering, and know when a sentence is finished so it can move on.

How it works

The video stream is read in fixed-length chunks that collect in an input buffer. At every step the translation model re-translates everything in the buffer, producing a draft of the current sentence that can still change. Training encourages each draft to keep the start of the previous one, so captions stay stable while they improve.

When the end of a sentence is stable in the draft, sentence commitment locks that translation into the history shown to the viewer and clears the matching video from the buffer. The buffer stays short, so the system can run on video of any length.

The model is trained the way it is used: on partial inputs, on pauses with no signing, and with several sentences of context.

Chunks enter a dynamic input buffer, the translation model re-translates the buffer each step, and a stable sentence ending triggers commitment to the history.
Chunks enter a dynamic input buffer, the translation model re-translates the buffer each step, and a stable sentence ending triggers commitment to the history.

Results

On sentence-level benchmarks (CSL-Daily for Chinese Sign Language and Phoenix-2014-T for German Sign Language), ReSCUE reaches higher BLEU at lower average lagging than SimulSLT and CTL++.

On long, unsegmented video it is compared with Uni-Sign, the strongest offline system, and with an oracle version of Uni-Sign that is given the true sentence boundaries, which a live system never has. ReSCUE comes close to the oracle's quality while cutting the delay from tens of seconds to about 1.6 seconds.

Long-form translation. AL is average lagging in milliseconds (lower is better); BLEU and ROUGE measure agreement with reference translations (higher is better). O-R uses ground-truth sentence boundaries.
MethodALBLEUROUGE
CSL-Story, Chinese Sign Language
Uni-Sign37,7222.1032.42
Uni-Sign (O-R)37,72224.8046.71
ReSCUE1,58823.6441.38
How2Sign-Long, American Sign Language
Uni-Sign53,6870.5124.16
Uni-Sign (O-R)53,68715.5939.61
ReSCUE1,58014.7138.42

My part

I wrote the proposal for a real-time system that translates a continuous video stream, the problem ReSCUE addresses. [Add what you built or ran after the proposal became the project: an experiment, a module, the demo.]