I have not used Collab. Within SE there are options to use CPP or Faster Whisper etc. I found Faster Whisper to be faster than CPP. For me what speeded up the transcription was getting a good GPU after I installed RTX3060 with 12gb it goes much faster but that's an expensive upgrade. Some people said 8gb is enough. But you need a card that uses CUDA. Also recently I discovered a parameter for Faster Whisper called -batched which became available from r239.1, it does batched inference and reduced the subtitle creation by 50%. Of course that may introduce more hallucinations.