How can I achieve this kind of waveform?

Hello everyone, I’ve recently run into a problem. How can I achieve a waveform like the one in the picture in Nuendo? You can see that the maximum level of the upper part of the waveform is -6, while the lower part is messy. The loudness of this audio is -16 LU.

My client requires my output to have a level of -6 and a loudness of -16. The entire waveform must be very flat, with no quiet sections and no loud sections. I’ve tried many methods but haven’t been able to achieve this. This is especially difficult because some voice actors deliver audio with extremely wide dynamic range, where the difference between the loudest peak and the quietest part is several tens of decibels. I’m really struggling with this. What I need now is a way to quickly control the level of all audio to -6 and the loudness to -16.

You could use the Nuendo Raiser Plug.In for this purpose:
Set the ceiling to -6db and then adjust the Gain to reach the desired loudness, check the integrated loudness in Supervison Loudness-meter:

As an alternative you could try those settings in the Export Audio Mixdown window:
this is actually pretty straight forward

Which gives pretty simular results:

Hi, thank you very much for your reply. I’ve tried the methods you mentioned, but the results are limited. For example, the level reaches the target in loud sections, but it doesn’t in quiet sections. If I boost the quiet sections to reach -6, the loudness will exceed the limit. I’ve also tried the ‘Integrated Loudness’ option when exporting, but in most cases, I can only choose one of the two — either loudness or level.

Maybe talk to the client what kind of loudness exactly, I don’t know what the standards would be here. Integrated loudness would be higher with lots of talk and no pauses, compared to talk with lots of pauses between short phrases, but the perceived loudness of the speech would be the same?

is it really necessary to have all phrases reach -6 peak-level?

I believe what the client is asking is that the maximum peak level is -6dbfs, not that every word must reach -6, and specially not every part of it, as it really doesn’t make sense to ask for that.

And even if that actually what they are asking for, I’d try to argue otherwise.

:pleading_face: I have communicated with him multiple times, but every time he just says, “I understand, I understand,” yet when it comes time to redo and revise, he says, “No, no, this doesn’t look smooth, the dynamics are too large, the quiet parts are too low, I need it to be very smooth.” So I want to ask everyone if there is a way, like in the image I mentioned in my question, to make the waveform align flat at the top while leaving the bottom as is. That way, at least the level would look stable. For example, like this.

I guess you could introduce a DC offset of 6dB with an audio editor (not sure if WaveLab Go allows to do that), then save the file with an integer bit resolution. Load it back into the editor and remove the DC offset.
This seems a very odd request to me.

This is a very unreasonable request, but he’s my client, so there’s nothing I can do.

What kind of source material is this? Voice over, set dialogue? Either case: clip gain is your first stop to roughly raise the levels per word or phrase (soft onss). Then use Raiser to maximise output level. Or use a soft clipper to take out transients.

But honestly this feels like a client you don’t need in the long run. Your client looks at a waveform and talks about smoothness, but do they listen to the sound?

Adding to what @klfnk2020 said:
Clip gain is always a first step if you are looking for decent results - depending on the source material. There are tools like Wave’s Vocal Rider which automatically ride the gain for you. It can’t compete with manual gain riding but it’s probably good enough to get you in the ballpark region. I am not so sure about Raiser. If I were you I would check out other options like single compressors as well.

However, after riding the gain you can also apply most of the above mentioned methods which should yield better results now. Without knowing what kind of source material we are talking about it’s difficult to recommend a specific option.

They listen to the audio, and they listen very, very carefully. Even the slightest noise and high-frequency hissing sounds can be picked out and sent back for revision (not sibilance). It’s more like a vibrating sound at high frequencies (similar to a very subtle vocal cord vibration). I’m starting to wonder if my microphone is broken… Most of the time, I can’t even hear those noises. I suspect they turn their computer volume up extremely high to catch every tiny flaw in the sound. Sometimes, I can’t even see the noise in SpectraLayers using the spectrogram. But they can hear it. They also compare audio files in Premiere. For example, in the image, as long as the waveform is very flat, that’s acceptable. If the fluctuations are even slightly too large, it’s not acceptable.

Yes, there is.
Since that would be a bad job on the engineering side of things you should opt to not to do it.
Your purpose as an audio professional is to do a good job.

What he requests implies a severe DC offset wich is absolutely wrong from an engineering stand point and presents no advantage at all.
In anything! It’s just a plain mistake.

In the end, if anyone else looks at your past work and sees that you have done that, it is bad for you and not for the client.
Bad engineering decisions are always on the hands of the engineer, not the client’s.

Yes, I could stand my ground, but he keeps asking for the same kinds of revisions. I’m really tired of it, but they pay me. I don’t want to lose the client. So I’m hoping there are some more experienced pros on the forum who can give me a way to quickly hit both precise loudness and precise level targets.

You have to use measurement tools in a program that includes measurement tools.
In Steinberg you have WaveLab that will allow you to see the LU levels in real time while you adjust the compression or limiting settings.
It also includes presets that can be applied to a waveform in order to achieve the desired maximum peak level with a desired LU (or LUFS) level.

The other option would be using plugins and for that i would recommend Junger Audio’s “Level Magic”. It’s a very complete Compressor/Expander/Limiter designed for broadcast but that can be used with superb results for editing or mastering.
There you can also set maximum peak level and adjust for certain LU levels, with real time control.
You can save presets with tests and you can apply different settings to the Left and Right channels.

Hope this helps.

Thank you very much for your help, Peace.

Hang on though:

I think that image (above) shows amplitude starting with zero at the bottom and positive values upward, right? The image you posted earlier,


shows both positive and negative values with zero (silence) being in the middle.

When you say that you want “to make the waveform align flat at the top while leaving the bottom as is” it seems to me that you’re misunderstanding things.

What the client seems to want is just a signal with a very limited dynamic range, nothing outrageous (unless I’m misunderstanding something).

If what you’re working with is what you showed earlier:

then it’s just the basics that are needed, nothing fancy. You can separate the middle section into their own events and raise the level of them so all events are roughly on par with each other. You can then either use clip/event-gain on the events or you can use volume automation to ‘counter’ the direction of the waveform. I normally do this first and then put dynamics after to get a consistent behavior from the processors: first a compressor with fast attack/release, then another compressor with longer attack/release times, and at the end of the chain a brickwall limiter at the required value.

You can use automated tools to level off if you don’t do volume automation manually, normally, like Reco29 wrote.

Hi abl,

I am wondering why no one here has hit on this yet:

You need to use compression and possibly limiting to achieve these goals quicker.

After that, you need to see if you have to adjust any passages that are still not loud enough / compressed enough.

Did your clients give you any audio samples to listen to?

Background: I once did the entire AT&T voice prompts for their new automated system, something like 800+ audio files, in a very weird proprietary format, similar to MP3, meaning small, but also bandwidth limited.

I spent a week recording all the prompts with the talent, then spend an entire day testing different methods until we got the best results after converting the audio into their proprietary format. It was very much touch and go, however once we had the exact method, it was easy as we had been saving presets as we went.

You’ll have to do something similar here.

Do you know anything about speech recording, editing and processing?

Compression and limiting (those are done with plugins) ?

I have also done a ton of audio books, and you do need to come up with a similar automated process for each voice talent, since processing hours of work by hand is a huge waste of time.

Always save presets as you go, and once you find the correct ones, write them down in a notebook, as well as the order in which you need to do them in to process all your audio files.

Good luck!

As mentioned above, this looks like heavy compression/normalization/limiting.

You can checkout Noiseworks Voiceassist. It lets you use ARA to level out spoken word very easily, just set the threshold that you want. I found it cleaner than using heavy compression/limiting.

accentize has a free dxLevel plugin.. might be useful