YuzuLingo vs video-to-text-japanese (katkaypettitt)
YuzuLingo stays in the browser. You cover burned-in subtitles on the video you are already watching, including YouTube, Bilibili, and other web players. There is no notebook to run and no video file to start from. Turn on Interact to read the burned-in line on your device, with furigana or pinyin when the language has them. Show that text always, on hover, or while you hold a hotkey.
video-to-text-japanese is katkaypettitt’s pair of Jupyter notebooks, in the repo video-to-text-japanese. You give it a local video and it transcribes the Japanese speech with IBM Watson Speech to Text. Light cleanup adds punctuation and removes filler words. One notebook is for videos longer than about 15 minutes, and one is for shorter files. You need an IBM Cloud account, an API key, and ffmpeg. They say Watson includes 500 free minutes a month. The sheet describes a pipeline that reads video frames and extracts spoken and visual Japanese. The README does not describe OCR of the picture. The README is dated April 2021. The last commit was about 3 years ago, and there is no release. MIT license.
Where each one fits
| Feature | YuzuLingo | Video to Text (Japanese) |
|---|---|---|
| Install in the browser, no notebook to run | ||
| Cover burned-in subs while a web video plays | ||
| No video file to start from | ||
| Read that burned-in line while the video plays, on your device | ||
| Furigana or pinyin on that line | ||
| Show that text always, on hover, or while you hold a hotkey | ||
| Transcribe spoken Japanese from a video file | ||
| Speech-to-text stays on the device | ||
| Add the line to an Anki deck |