SLP 13 Unmix Two Voices - Bugreport

Ok, just bought the update to SLP13, looking forward to the new Unmix Two Voices - as that would tremendously improve my workflow with reporter style interview audio.

@Robin_Lobel Alas it has bugs out of the box.

Interview with a male reporter and a male person.
Source is a 32bit mono WAV rendered out from Davinci Resolve.

Not only do the voices get mixed up sometimes on the first layer, sometimes on the second layer (even mid sentence). There are also repeating parts where there is a hole in the upper frequencies above 3.8 kHz - completely erased. So much on non-destructive unmixing…
It seems the hole is following a pattern - every 30 seconds comes a hole for 4 seconds.

I can provide audio files to test this in a private link.

Repeatedly loss of upper frequencies, leaving holes.

The marked region is all the same speaker voice, but chopped up to layers by Unmix.

System specs:

  • SpectraLayers 13 Pro 13.0.0

  • Windows 10 Pro 22H2

  • AMD Ryzen 9 5950X (16 cores)

  • Gigabyte B550 Vision D-P

  • Corsair Vengeance LPX 128 GB (4 x 32 GB) DDR4 3600 MHz C18

  • PNY NVIDIA RTX A6000 (Ampere) 48 GB

  • NVIDIA WHQL Studio Driver 553.50

  • Seasonic PRIME TX-1000 80PLUS Titanium 1000 W

  • WD_BLACK SN850 2 TB NVMe PCIe Gen4 (system)

  • Crucial P5 2 TB (fast cache)

  • 4 x Crucial MX500 4 TB as 16 TB RAID 0 (big cache)

  • LSI 8-port RAID controller

  • 8 x 16 TB Toshiba Enterprise MG08ACA16TE as 104 TB RAID 5 (work files)

Yup I can confirm similar results here.

Just like Unmix Multiple Voices never worked for me, these new modules seem to follow the same path, unfortunately.

I tested further with other interviews. It doesn’t matter if it’s a male and female voice or not, or very different sounding voices - the module just fails 100% all the time.
All professional interviews with no one talking over each other.

I mean - how did they beta test this? This is not a bug where you need to drive the tool into an edge case to get it shown…

That’s precisely what I get, and what the previous tool gave me, too.

Even when there are two radically different voices without overlapping, It always jumbles everything up, like attributing only the end of a word to a different voice, things like that.

I don’t get it either… =(

Thanks for this guys,

Robert, fantastic presentation as always. I wish the results were better :-/

Same here doesn’t work at all, even with 2 very different voices

FYI, not seeing that here (Win 11). Cannot repro the bug. Do you get the same result with 24-bit 48kHz Wav stereo audio files?

Tested with 24-bit 48kHz WAV mono and stereo - unfortunately always the same results.

@Robert_Niessner Thanks, the 4khz gap is now fixed here and will be in patch 1.

@henrique_staino @ctreitzell @gugu

However this module is not meant to unmix long interviews or podcast, it’s primarily designed to unmix people talking over each over, as described in the documentation and the demo video. Interviews or podcasts are for the Unmix Multiple Voices module.

In your example I don’t see any crosstalk at all, or even two speakers in the first 30 seconds, hence why it’s producing those unpredictable results.

There was an internal debate with this feature to show a warning message if users were going to attempt using it on audio samples longer than 30 seconds (it’s unlikely overlapping voices last more than 30 seconds). It was decided against it, but maybe displaying such a warning should probably be better to avoid such abuse of the module ?

Hey @Robin_Lobel that’s nice to know! I do remember using it like you described, but it was with two women talking and they had somewhat similar voices.
I’ll give it another try with that particularity in mind when the opportunity arises in my work here.

But since you mentioned it, I must say, Unmix Multiple Voices has never worked for me no matter what I tried it on. In the end I stopped trying because it was always completely off and more trouble than doing what was possible by hand.

I think the warning would be welcome, yes! =)

Ok, but sorry that module was the main reason for me to update as it would have helped me a lot and save time. I watched the presentation video and not once I got the impression that this would not work as intended.

Quoting Steinbergs description of the feature:

Unmix Two Voices

Automatically separate two voices without having to register voice profiles. This is one of the workflow speed and efficiency improvements which are a hallmark of the new edition.

If that isn’t misleading I don’t know what misleading is.

In the second screenshot I have clearly marked where the second speaker is talking - this is a typical interview situation where the journalist is asking a question and a person is answering.

EDIT: So - when will get a working Unmix Multiple Voices then?
I’ve tested Unmix Multiple Voices in SLP13 now - seems like this is working much better now. At least with the several recordings I tested it here.

EDIT 2:
Tested now another more difficult audio track with less ideal audio from a satirical piece - a woman and a man acting: Even with exact voice registration Unmix Multiple Voices fails spectacularly as always.

My main gripe is that cannot be automated which kind defeats its purpose for me - hence I was so eager for the new Unmix Two Voices module.

In my workflows I am looking for the fasted way with the most automation to get things going:

I have 40 video interviews each about 7 minutes long.

  • Export each audio track from Resolve.
  • Auto-batch those tracks in a prep setup through RX
  • auto batch them through Unmix Noisy Speech in SLP12.
  • load speech and noise layer then back into SLP12 to check for false positives like f and s sounds in the noise layer, put them back into speech, save that
  • master speech in RX
  • manually separate speakers in SLP12 for transcription
  • push audio through AI transcription (diarization there has never worked correctly)
  • SRTs and preview videos (with combined speakers) get fed into a custom build Interview Assist tool for the journalist to work with.

Every manual work step I can avoid is a huge time saver.

Currently Unmix Multiple Voices needs me to register all the voices then separates into voice and non-voice layers (which are mostly breaths). I then have to check again for false positives, put the breaths back for each voice, delete the non-voice layer and export the voices layer.
Then I have to deal with the naming conventions so nothing breaks in the rest of the workflow chain.

@Robert_Niessner thanks so much! Your in-depth logging/ reporting is deeply appreciated by me; teaching us all!

When you refer to 2 voices are you talking about a lead vocal and a chorus on separate tracks to be stemmed?

No, I am talking about interviews with talking heads.

no idea what that means-please educate me

Maybe on the road to nowhere?

I am filming interviews with people, where a journalist is asking them questions.
The journalist is holding a microphone in his hand, says his question into the mic then tilts the mic over to the person giving the answer.
This means both speakers are recorded onto the same audio track.
What I need in post is automatic separation of both voices into layers.
Currently I am doing this manually but this could clearly be automated in the future.

got it

Hey @Robin_Lobel coming back here just to report that unfortunately I have been very unsuccessful in using Voice Decrosstalk even following your instructions.

On very clear, short overlaps, where the difference in volume is large, where one of the voices is starts well before the other, and is well longer, it simply does not do what it says on the tin no matter what I do. And in the very rare occasions it has worked, they are just so rare that they seem more like happy accidents.

Undoubtedly I’ll have to abandon trying to use this module untill a better version comes out, just as I’ve had to do with Unmix Multiple Voices, Unmix Two Voices, Declick and DeHum.

I hope we can find a way to make them work as they would be a nice addition to the other very useful modules.

Thanks!

@henrique_staino can you share a sample on divideconcept [at] gmail dot com ?