Thanks for subs. Is there an easy way to filter out the Chinese text from the subs? And how to use the ignore file? Thanks in advance.
Attached: ADZ-269, FAX-133, IPTD-534, JUFD-182, ONED-590, ONED-939, ONSD-724, ONSD-724, RBD-369, ROE-009, ROE-010, ROE-014, ROE-015, ROE-018, ROE-019, ROE-020, ROE-021, ROE-023, ROE-024, ROE-025, ROE-027, ROE-028, ROE-031, ROE-032, ROE-033, ROE-035, ROE-036, ROE-043, ROE-053, ROE-056, ROE-066, ROE-085, ROE-103, ROE-105, ROE-111, ROE-112, ROE-114, ROE-116, ROE-122, UMSO-129, URE-006, URE-013, URE-030
Yeh sorry about that. I like to see the Japanese and English to try and catch bad transcriptions/translations. So I usually watch with dual subs. I went ahead and wrote a script to just extract the english out. I've attached some example subtitle files from the machine translations. Most of these havent been touched, so there are a lot of bad transcriptions/translations in them.
As far as the filter text file. That is for anyone using Whisper (such as on Google Colab). After transcription simply compare transcribed lines against the list and filter out the bad lines (e.g. hallucinations). I've been experimenting with using a VAD, and then automating post-processing/cleaning. I just need to figure out a place to upload all of these. I think I've done ~1000-2000 movies so far... I don't know what the file size limit is when uploading rar files of srts.




















