Skip to content

Comment on Nerd-dictation, hackable speech to text on Linux

Comments

You know... I have an idea. How about we use vosk and this tech to integrate with ffmpeg somehow so that peertube videos can get subtitles while being transcoded. Once we get English SRT, we could use libretranslate to translate that English SRT to multiple languages.

This could be similar to what YouTube does with it's automatic subtitles. What do you guys say?

The package that this project is built on (vosk-api) mentions & includes some examples to demonstrate exactly that type of use case with ffmpeg:

"...continuous large vocabulary transcription, zero-latency response with streaming API, reconfigurable vocabulary and speaker identification ... can also create subtitles for movies, transcription for lectures and interviews."

* https://github.com/alphacep/vosk-api/blob/master/python/exam...

* https://github.com/alphacep/vosk-api/blob/master/python/exam...

* https://github.com/alphacep/vosk-api/blob/master/python/exam...

(Edit: Also, thanks for introducing me to "libretranslate", looks like an interesting project.)

great. someone should link peertube github with this. i am sure the great people will do it much faster and more elegantly :-)

I just checked and apparently they are already aware: https://github.com/Chocobozzz/PeerTube/issues/3325#issuecomm... :)

(Tho I'll admit I have no idea what "bluffing" means in that context. :D )

cool.. https://gitlab.mim-libre.fr/extensions-peertube/plugin-trans...

so this already exists. damn i thought i was the first one to think of this. still, this existing means there are people who think faster than me

Sounds great! Do it!

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.