NeurIPS 2026
ReSCUE
Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation
Live captions for sign language that keep up with a signer for minutes at a time, instead of one pre-cut sentence at a time.
- Authors
- Sihan Ren*, Gaozheng Li*, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye Wu. *Equal contribution.
- Institutions
- ShanghaiTech University, University of Derby and DGene. Work done during an internship at DGene.
- My part
- Wrote the proposal for a real-time system that translates a continuous video stream.
The problem
Real-time sign language translation could let a deaf signer and a hearing listener talk without waiting for an interpreter. Most translation systems, though, work offline and one sentence at a time, and they expect someone to have already cut the video into sentences.
Real signing does not arrive that way. It runs on for minutes, with pauses that are not sentence breaks. A live system has to decide what to show before a sentence is over, revise it without the caption flickering, and know when a sentence is finished so it can move on.
How it works
The video stream is read in fixed-length chunks that collect in an input buffer. At every step the translation model re-translates everything in the buffer, producing a draft of the current sentence that can still change. Training encourages each draft to keep the start of the previous one, so captions stay stable while they improve.
When the end of a sentence is stable in the draft, sentence commitment locks that translation into the history shown to the viewer and clears the matching video from the buffer. The buffer stays short, so the system can run on video of any length.
The model is trained the way it is used: on partial inputs, on pauses with no signing, and with several sentences of context.

Results
On sentence-level benchmarks (CSL-Daily for Chinese Sign Language and Phoenix-2014-T for German Sign Language), ReSCUE reaches higher BLEU at lower average lagging than SimulSLT and CTL++.
On long, unsegmented video it is compared with Uni-Sign, the strongest offline system, and with an oracle version of Uni-Sign that is given the true sentence boundaries, which a live system never has. ReSCUE comes close to the oracle's quality while cutting the delay from tens of seconds to about 1.6 seconds.
| Method | AL | BLEU | ROUGE |
|---|---|---|---|
| CSL-Story, Chinese Sign Language | |||
| Uni-Sign | 37,722 | 2.10 | 32.42 |
| Uni-Sign (O-R) | 37,722 | 24.80 | 46.71 |
| ReSCUE | 1,588 | 23.64 | 41.38 |
| How2Sign-Long, American Sign Language | |||
| Uni-Sign | 53,687 | 0.51 | 24.16 |
| Uni-Sign (O-R) | 53,687 | 15.59 | 39.61 |
| ReSCUE | 1,580 | 14.71 | 38.42 |
My part
I wrote the proposal for a real-time system that translates a continuous video stream, the problem ReSCUE addresses. [Add what you built or ran after the proposal became the project: an experiment, a module, the demo.]