Automatically tag non-speech audio
under consideration
Gabe Michalski
Merged in a post:
Laughter Detection/Transcription/Tagging
Matt Hayward
Laughter in a recording is not transcribed, and shows up as "silence."
Thus bulk remove of silence removes laughter, and other non-word communication (gasps, sighs, etc.).
Ideally transcription would tag common exclamations as Laughter, Sigh, etc.
Barring that, it would be nice if it could distinguish between actual silence, and non-word sounds in transcription so that shorten silences wouldn't cut out things like that, or could be at least be configured not to.
It seems like the "tagging" functionality would be a good match for this, however that does not work on Multitrack, and this feature would be needed in Multitrack as well.
Lazar
Couldn't agree with this more. The remove silences feature, effectively, is a laughter identifier, and that actually is way more valuable, but it really does need to differentiate between the two better.
Autopilot
Merged in a post:
The one thing no one in the AI editing space does: non-speech sounds in the transcript
D
David Etler
Tagging overtalking and laughter can be used as signals to Overlord of important moments, helpful in finding clips to create a task Overlord isn’t doing that well right now. And rather than treating these as a problem by removing them as “pauses” the transcription could allow us the option to leave them alone, as conversation features that add vital context to the transcript for editors. Also helpful for accessibility purposes! HINT: I don’t think this is something anyone else is doing, which means it’s a differentiator if it can be done…!
D
David Etler
Totally. I’d add overtalking to this. Tagging overtalking and laughter can also be used as signals to Overlord of important moments, helpful in finding clips to create.
J
J D
Perhaps these can be handled with a script layer?
I'm thinking that a Composition might be able to be processed through several AI filters to find interesting information.
We see already how amazing https://notebooklm.google.com is able to handle audio, but imagine if you can run three filters on each composition:
1) Sound and music - creating 1 script that can be searched for coughs, laughter, singing, etc.
2) Transcript - as usual but faster, like the speed of notebooklm.com for mp3 files
3) Video - looking for people, objects, images, animation, clocks and timers, text, etc. and creating a script with those objects named in the timeline.
J
J D
Another advantage to doing it that way is that we can imagine adding 3 fancy captions to the result that would allow a blind person to listen to the visual activity or a deaf person to see the audio activity in addition to the transcript.
So consider this an accessibility function in addition to making search and AI generation of the content more interesting because clips will be easier to find if the details are in text and accessible by tools that use the LLMs.
S
Support Team
Merged in a post:
breath
M
Miko
Would be so great to have a big breath/gasp detection to add to the fillers. Thanks.
Gabrielle M.
This would be AMAZING to see in the software! So frustrating when it counts breaths as blank audio clips. >:(
Autopilot
Merged in a post:
Recognize laughter
C
Carson Gale
I love the remove silences and trim content features, I just wish it wouldn't also remove the laughter. Would be great to not lose laughter by default when I do these mass quick edits.
D
David Nadelberg
This is great. Ideally, this feature would also be able to label recordings of crowd/audience laughter and other noises. If you are editing content that was recorded in front of a live audience-- group laughter, shouts, gasps are key elements that need to be listed in the script. So I am hoping that this tool will be able to recognize group laughter, not just the type of laughter heard in a quiet, 2-person conversation that lacks crowd noise.
S
Shannon Post
Supporting this feature request
Load More
→