Omnivocals Beta in Cubase 15 is great and is transformational for song writing without having the singer on hand all of the time. Having solved various phonetics and word cadence issues (sometimes by changing lyrics to something Omnivocals handles better!), there is then a lot of editing of automation for Attack, Power, Vibrato to arrive at something more human. I then render down to a wave file and put it through Melodyne to edit the pitch (sometimes to pull it off centre to be more human!) and particularly to try to remove the very digital pitch sweeps (think Autotune!) that can occur between notes/words. This then goes through some outboard analogue gear to introduce some harmonics etc. After all of this there is still a slightly “metallic” digital quality to the singing that I’m struggling to dial out. I’d be interested in how others are processing Omnivocals to get closer to a human voice? Thanks.
For the most part, I don’t bother, because the way I’m using Omnivocal is either meant for temporary purposes or for context that won’t be highly exposed in a mix. More on that below, but a few thoughts based on what you’ve indicated you’ve already done and just some thinking out loud first.
I just reviewed the notes I’d made in my original Omnivocal experiment (with the version included in the first Cubase 15 release), and I’m not even sure I automated the parameters you mentioned, though I did automate pitch bend, initially by just overdubbing, then editing to address some issues. I also turned the formant control slightly lower than default as I felt it worked better with both voices. If you’d like to check my notes on that experiment (and there are also links to the results for both singers), here is that post:
Please note that the title of the thread does not correspond to my view.
(I was just replying, actually quite a bit after the initial post.)
In that case, I was doing the minimal to do a quick experiment, with the processing I used on Omnivocal being putting it in an environment (via the UAD Ocean Way plugin) and using a preset in a vocal chain-type plugin (Waves CLA Vocals – I’d probably have been more likely to use UAD Topline Vocals if I did that today, but I didn’t have it at the time, and I got a better “instant” result with the CLA plugin than with Cubase’s VocalChain presets).
In general, when I do actually use Omnivocal in a recording (and I’ve only done this once to date), I process it similarly to how I’d process my own vocals in the same context. And that context has only been layering Omnivocal voices with my own background vocals, to try to get more of a group vocal feel, as opposed to one that is only one singer overdubbing all the BGVs. I’m not needing to do anything in Melodyne because my own vocals are layered with the parts. Also, the parts are buried enough in the mix, that any minor issues on the “more human” front won’t be noticeable. (I also blended my vocals higher than the Omnivocal doubles.) I’ve since switched to using IK’s ReSing for that function, as it’s way easier to be able to just duplicate my own vocals for processing with ReSing than it is to try and recreate MIDI performances of what I was singing to use with Omnivocal.
My main use of Omnivocal, though, has been for a placeholder lead vocal while working on the instrumental arrangement for a recording. It will be replaced by my own lead vocal later, but using Omnivocal to substitute avoids needing to set up a mic, worry about tempo changes with an audio recording, changing my mind on lyrics (if I’m working on the recording in parallel with writing the lyrics), etc. And I find having a “vocal” with lyrics quite helpful compared to my past technique of just using some vocal “ooh” patch on a synth or sampler. In this case, I really don’t worry about it sounding more human – it is close enough for these purposes – or its mispronunciations (though I will at least correct any that just aren’t working in the context of my use of a word – e.g. sometimes we want “fire” to be one syllable, and other times we want it to be two, so phonemes are necessary in the former). I just run the Omnivocal track through UAD Topline Vocal in that case to get it to sit better in whatever stage of the arrangement I’m at.
There’s also one new thing I’ve tried on my current project-in-progress. I rendered the Omnivocal part to audio, then I used ReSing to change the vocal to another voice that was more compatible with my own vocal tone. That was actually a major improvement on the “humanity” side. It does, of course, mean that, if I decide to change tempo or automate tempo changes, I’ll need to re-render and re-process the part (or just shift back to using Omnivocal itself) if I’m not ready to just record my own vocals at that point.
One more thought on the note you made about the “metallic” aspect: I wonder if using a plugin that detects and reduces resonances could be useful. I’ve used Waves Curves Equator or that general function, but I know there are other plugins that do similar things. Might be worth a try if you have it (I think there’s a time-limited demo available if you don’t).
Thanks Rick, useful information and suggestions. I’ll check out the Waves plugin and report back!
One thing i noticed from putting a render of Omnivocals through Melodyne is that Omnivocals is clearly using an algorithm based on listening to alot of real singing (as one would expect for an AI instrument). For example, it often mimics the lower pitch at the start of a word before hitting the note that one writes in the midi track. This usually sounds ok, but can produce a slightly unnatural slur. It does also tend to generate quite slow pitch transitions between notes, which can create that Autotune sound. In Melodyne it’s simple to speed up these transitions to reduce that effect and experiment with setting the correct pitch straight away at the start of words. Oh, and it does sometimes overdo the breaths at the start of phrases or in places where a real singer probably wouldn’t breath so audibly. Again, these can be reduced easily in Melodyne. Of course if one is just using Omni to provide a vocal template, then no need to worry about any of this!
Note that Omnivocal isn’t an AI instrument. It is vocal synthesis. Yamaha has a long history on that front with their Vocaloid technology and product line – I reviewed one product based on that (Zero-G’s MIRIAM Virtual Vocalist) for the now long-defunct CakewalkNet ezine way back in October of 2004 – way before any AI technology was making its way to products.
I wonder if working with the Attack setting might help here? I have noticed that it often can’t deal with short notes very well, so I sometimes extend them to allow it to get in final consonants instead of dropping them in the transition between words.
I’ve noticed this occasionally, too, but, then again, sometimes real singers do this, and that is a simple thing I deal with all the time with my own vocals in the process of editing fades and such after comping. There are also plugins (e.g. Waves DeBreath) that can be used to reduce overzealous breath sounds, be it at the extreme of removing them altogether or just reducing them to fit tastes and feel more natural. However, I do prefer dealing with that manually since cut points in comping always at least occur in the vicinity of breaths (though potentially in other places as well), so my normal cleaning up after comping will already be dealing with fades or crossfades in those areas.
As for Melodyne, I know it is super capable, but, at least to date, I haven’t used it for timing-related changes. I really only use it for tuning, with only extremely occasional forays into other areas (e.g. formant tweaks and sibilance reduction). Some of Celemony’s videos have occasionally tempted me to consider using it in more advanced ways, but I tend to have other tools and/or methods that I’m more familiar with.