Mixing Workflow in Cubase – Looking for a Structured Approach

Hi everyone,

After studying various tutorials, videos, and articles on mixing, I’ve put together a general mixing workflow for Cubase. I understand that mixing is a creative process with no single “correct” method, and the ultimate goal is simply to make the song sound good.

However, as a beginner, I am looking for a practical, non-genre-specific workflow that can serve as a solid foundation. I would appreciate feedback on the following approaches.

Situation 1: Mixing Client-Provided Audio Tracks

Let’s assume a client provides:

  • 50 audio tracks (lead vocals, harmonies, guitars, drums, etc.)
  • A reference track (mixed but not mastered)

Reference Track Question

Do reference tracks need to be mastered before they are useful for mixing purposes, or can a mixed-only track serve as an effective reference?

My Current Workflow

  1. Create a new project and set: Tempo, Sample rate, Bit depth and Buffer size (typically 512–1024)
  2. Import all audio tracks and the reference track
  3. Set up the Control Room for A/B referencing
  4. Organize the session: Rename tracks, apply color coding, create folders and create Group Buses
  5. Keep all channel faders at unity gain initially

Gain Staging

Using the Pre-gain control in the Mix Console, I reduce the level of tracks that are clipping. My typical targets are:

  • Peak level between -10 and -8 dB
  • RMS between -18 dB and -12 dB
  • At least 6 dB of headroom

After gain staging the individual tracks, I often notice that some Group Buses and even the Stereo Out still clip.

  1. I have read that this is caused by summing. My understanding is that multiple signals combine and increase overall level. Could someone explain the technical reason behind this more clearly?
  2. Should I gain stage the Group Buses at this point?
  3. Once gain staging is complete, I usually pull down all faders and build a static mix from scratch, starting with the most important elements (typically drums and vocals). Is this a good approach?

Reference Track Level Matching

I frequently hear that the reference track should be level-matched to the mix.

  1. At which stage should level matching happen?
    1. Before gain staging?
    2. Or before adjusting volume faders?
  2. If the reference track is louder than my mix, I lower its level using pre-gain or the channel fader. Is this the recommended method?
  3. If the reference track is quieter than my mix (as is the case here, since the reference track is not mastered), should I:
    1. Increase the reference track pre- gain levels?
    2. Or push its fader above unity?

Processing Stage

After completing the static mix

  1. Apply panning
  2. Use stock EQ in the Mix Console for cleaning unwanted frequencies and frequency separation
  3. Apply compression plugin: On individual tracks when needed and on Group Buses for glue and control
  4. Create FX Channels and add: Reverb and Delay plugin (I generally use send effects rather than insert effects.
  5. Apply automation as the final step before the mix is completed.
  6. Does this sequence make sense?

Situation 2: Mixing My Own Production

In this scenario, my project contains approximately 45 tracks, including Audio tracks, Instrument tracks, Multi-timbral MIDI tracks and Loops

Bouncing vs. Mixing Directly

Before starting the mix, should I:

Option A: Bounce all Instrument and MIDI tracks to audio and then mix entirely in Audio?

Option B: Keep the Instrument and MIDI tracks active and mix using their corresponding audio return channels?

My Thoughts

Advantages of bouncing to audio: Reduced CPU load, Improved session stability, consistent playback

Disadvantages: Loss of MIDI editing flexibility.

I have occasionally experienced erratic VST behaviour when mixing directly from virtual instruments. For example: timing/groove inconsistencies, notes triggering incorrectly and sounds playing out of place.

This seems to occur especially when jumping between different sections of a song, whereas playback is stable when starting from the beginning.

Has anyone experienced similar issues, and would freezing or bouncing tracks be the preferred solution?

Mixing in mono versus mixing in stereo is another topic that is often discussed. However, I am still trying to fully understand the purpose and practical application of mono mixing, particularly for tasks such as setting volume balance, identifying phase issues, and addressing frequency masking. As a result, I have not yet incorporated mono mixing into my workflow. So far, I have mixed primarily in stereo, as most real-world listening environments—headphones, earbuds, car audio systems, and home speakers—are stereo.

Thanks in advance!

Well, the explanation lies in the word “summing” :slight_smile:. If you sum two numbers, you usually get a bigger number. Now with audio, this is a bit different as we have positive and negative values, so of course also the opposite can happen (e.g. if two signals cancel out), but if at any given point in time, you have two samples that are a positive value, they get summed up and as a consequence, the resulting signal is louder.

Personally, I rarely get clipping on busses when summing tracks that have been roughly gain staged with the method you described, but if it happens (maybe lots of tracks going to one bus), you can either lower the source tracks further or lower the input on the group. Or you could just wait and see until you have finished your static mix and see what is left…

Sure, if it works for you. I do pretty much the same, except I don’t pull the faders down but simply solo the tracks I begin with. But whatever works for you.

Overall, I think you already have a very good and structured approach. Could it be that you might overthinking things a little bit? Nowadays there is a lot of “pro advice” floating round the forums, Youtube and such, but it doesn’t mean that it is a) always good advice and b) necessarily useful for you. Each person’s workflow is different. For example, if you tried mixing in mono but didn’t find it helpful, just don’t use it.

I am really not a good mixer, but I think I have learned over time to not take all those “rules and regulations” about mixing that floating around the net too strictly. I try them, if I think they are useful I incorporate them in my workflow without being too pedantic about it (e.g. with gain staging I am quite happy if things are roughly ok, I don’t even use a VU/RMS meter anymore, just the channel meters)…

What exactly do you mean by “reference track”? Because usually that term is used for one or many tracks that are already mixed and mastered that you think sound good and that you use to compare your mix to or use for tuning your monitoring. In that case, level matching helps imho, and there are plugins for that (e.g. Metric AB)

Or do you mean a rough mix that is sent to you by the client? But wouldn’t that only be used to get a feel for how the client thinks the mix should sound? You are not going to rebuild their mix as identical as possible, aren’t you? In that case, why bother level matching that?

Yes, a reference track is a professionally produced track that sounds good and is used as a benchmark for comparing mix. In most cases, reference tracks are both mixed and mastered.

However, in my case, I am using a reference track primarily to practice mixing and train my ears. I downloaded a set of multitrack audio files along with a reference track (a rough mix rather than a mastered version), and my goal is to recreate the rough mix as closely as possible. I wanted to match the volume level of my initial mix to that of the reference track.

My mixing workflow has evolved drastically since Folders with Groups have been introduced. I set up a new template using these as containers with my go to plugins and presets and my default mixing routing.

I am spending a lot less time now importing tracks and getting everything organized and ready for mixing : I just drag imported tracks into the proper folder and it gets colored, routed and with default plugins and sends. I then mix mostly using folder with groups instead of individual tracks.

I must say being able to switch easily between stereo or mono (groups or tracks) is also a workflow enhancer for the mix preparation step.

The idea is that you pick something that you think sounds great and try to get close to that. Doesn’t really matter in and by itself if it’s mastered or not if that’s the sound you like.

  1. If you add stuff, numbers get bigger. It’s literally that simple.
  2. You should decide where to adjust levels based on what your goal is and what the problem is. So you start by identifying that. If you have set up a bunch of group tracks and you have dynamics processing on them “for glue” for example, and your problem is that the master is clipping, then you can’t just lower all tracks to fix the problem with clipping without creating a different problem - because you would have lowered the level going into those group tracks’ dynamic processing. In other words if you have a compressor on a group and you need to lower the overall level you need to change the group track output, not the signal going into it. If you change the input the compressor will work differently and it will change the balance and sound of the mix.
    In this case, if your mix is “complicated”, the fastest solution may be to just pull down the master fader. What some people do is put a separate mix-group track before the output bus and route all tracks to that mix-group so they can adjust overall level on the mix-group. I don’t personally do that but I know it’s common.
  3. Whatever works for you. My approach was often to start with faders down and then I’d bring them up to get a balance going, but I would keep an eye on levels as I did that. In other words this step 3 you are mentioning basically included the original gain staging. Mind you that back in the days when I did music mixing the tracks I worked on were usually recorded well and already at a decent starting level.

I really think a lot of people overthink all of this and make it more complicated than it has to be. If I were to mix pure music today and I used a reference track my approach would probably be this:

First, set reference track fader and Control Room monitoring so that the output level is reasonable. Loud sections should sound loud, soft should sound soft. Obviously the reference track needs to be at a reasonable level before it hits the output path with peaks around 0dBFS.

Then, with all faders down just bring up instruments and make them hit desired loudness and balance, going by ear. As I do this I also work on tracks together and add processing as needed. Is the drum kit too dynamic? I’ll add compression to individual items and probably route to a group and put more compression there. As I do this I continuously adjust balance and check with reference.

Again, to me it’s a lot more dynamic. If you set static levels and you’ve done zero panning it means that you have already “mixed in mono”, effectively (assuming there is literally no panning and no stereo elements). But the whole point of stereo isn’t just that you get a wider field, it’s that you can fit more stuff using panning (in my opinion). This means that as you bring up your faders during your “static mix” you’re going to have to be conservative because things may clash with each other, but once you pan them the problem goes away. So why not pan while doing your “static” mix?
And to me the same applies to EQ and compression. When I start pulling up the faders I hear how some elements need to be shaped in order to work together, so I apply effects. I don’t necessarily always think in “steps”. Of course, this is easier with experience.

Adding to the idea that this is dynamic: if you need to make space for a lead electric guitar that plays a solo you might lower some elements in the mix that are in the same frequency range, and that would mean automation. Or, you might just pan them out of the way instead. Or, you might EQ them instead. Or you might do a combination of that. So as you go along you’ll always revise your approach dynamically based on what the song needs.

Sorry for the long post.

On gain-staging,

My personal preference is I do this at the event level to optimize waveform height and volume for any editing that still needs to happen, for example in Spectralayers. I try to gain events by 6db, as this is a perfect bit and has no affect on the audio, but sometimes it needs to be by 3db to get a nicer relative average to everything else.

With summed clipping, sources that share similar frequencies will have more of an affect on this. I pull all the faders down, and then bring up only the bass and drums and get them balanced - sometimes by ear, sometimes by using the RMS method:

Once that is done, I start to bring up the faders of other things and panning, usually vocals first (if there are vocals) and other primary instruments in some logical order, then auxiliary instruments and develop a rough balance.

The next part is just a personal preference and it’s material dependent on what type of mix I’m doing, but sometimes I then select all the faders with Q-Link/relative enabled and bring all the faders up until I start to see the clip light intermittently turn on in certain/loud/busy sections because it’s an indication of frequency build up or transients that need to be tamed and from there I start making decisions on compression vs gain editing, filtering, etc, etc… The clip light being an indicator of sections or even individual beats I need to be aware of.

One broad thing you could do to speed up your workflow, is have a Mixing template with your go-to inserts already loaded, and with the new Folder-Groups, this makes bus assignments fast as you can drag your drum tracks into the drum folder, and have them routed.

I don’t depend on reference tracks too much, and definitely not at the start of a project - I want to approach the beginning of a project as a process of exploration, discovery, and feeling, and experimentation - not as copying. I want to hone in on what I naturally feel is authentic, interesting, and exciting and develop the mix based on that - not what has already been done. And I do say ‘no’ to clients and explain this to them, and if they don’t like it they can go somewhere else.

I will use the reference track near finish just to check whether I am in the ballpark or not and things are properly balanced.

Mind you, I’m at a stage now where I am confident with my ability and sensibilities - if you’re just starting out, maybe use reference tracks throughout the process - but even then, I sort of encourage people not to too much because it’s important to develop your own approach as your own approach is what will yield your result, that stands out from others.

Mono is useful for the reasons you mentioned, also, for the reason that your ears/speakers/room are not perfectly balanced. So, if you pan your hi-hat right mostly, and your right ear has some high frequency deficiency - you might not be mixing your hi-hats accurately.

Apart from that, I wouldn’t worry about it too much as mixes start to sound pretty boring pretty quick when they are overly “mono compatible”.

Wow! This is an amazing piece of information.

Thank you for explaining all this in such great detail. A big relief to my inquisitive mind :slight_smile:

Thanks for sharing the links. I’ll check them out. I appreciate your time and valuable insights.

Hi! Are you saying that the Control Room Level knob and the Reference Fader should be adjusted so that their levels match?

I have a few questions about the Control Room:

1. What is the purpose of the Control Room Level knob (the large red knob that is set to 0 dB by default)? It seems to only increase or decrease the monitoring volume. Does it affect the mix output in any way? If not, what is its intended use?

2. What is the purpose of the Reference Level button? Whenever I click it, the Control Room Level switches between -20 dB and 0 dB. What exactly is happening, and what is the purpose of this function?

It’s for those who do not have a hardware volume controller to their speakers. It does not affect the mix.

I probably didn’t explain very well. I’ll try again later. I’m a bit tired right now. Gotta eat.

Like the previous poster wrote, it doesn’t affect the mix, and it is a substitute for / or complement to a physical hardware controller for speaker playback level. In my room for example I have Nuendo and a MOTU 16A interface, and my speakers go directly to the interface. Since I don’t have a separate hardware level controller I use Control Room in Nuendo for that purpose.

The way I see it that button is there so you can always return to a specific value.

For example, in my studio I have my signal chain set so that my reference level played back in Nuendo (on a source audio or group track) when played back through Control Room = about 79dB SPL at the listening position (I forget the exact number). Whatever the Control Room knob is at when I reached that calibration level is what I set as a reference level in preferences.

Now, if I’m mixing for like 10-14 hours with few breaks my ears are going to get tired if I’m doing it at the reference level all the time, so I might turn down that CR Level knob. Instead of guessing what my reference is I can just hit that Reference Level button and I’m back to reference.

This way, if something sounds “loud” in my space when that button/option is on then it is loud, because that’s how my room is calibrated. So I can always mix a bit softer and just get back to reference really quickly to get a feel for if it’s actually as loud as I think it is.

I hope that makes sense.

Really the whole idea is that you calibrate your playback system to a reference level and you have a way to get back to that level at the press of a button/key. Then you can mix by ear, and what sounds loud is loud, in the mix.

There’s a bit more to it than that, but that’s the gist of it.

I would add, it can also just be used to do 'quiet mix checks". ie you want to analyze your mix at a variety of levels, sometimes listening to a mix quietly reveals more than listening loud.

Certainly. You could use it ‘either way’.

Okay got it, so the volume knob on my Focusrite interface, which is connected to the Beyerdynamic DT 990 headphones, serves the same purpose as the Control Room Level knob.

Sorry, I’m still not completely clear on this. When you say “reference level,” are you referring to the level of the reference track itself, or is it something else?

How can I determine what the reference level is?

Also, how do I calibrate my playback system? I’m using a Focusrite USB interface and Beyerdynamic DT 990 headphones.

The level at your ears.

You determine it yourself based on your preferences and requirements.

Your Focusrite device is going to have a manual that talks about setting it up. You should start by reading that. You can link to the manual here and other people can help if you need it.

In my case I just play back a reference mix that is on an audio track at unity, and then starting with the Control Room level all the way down I turn it up gradually until it sounds appropriate. That level is what I then make the reference level in preferences.

Actually, Asunder, I don’t know them. Please tell me.

Cubasodus 20:

1 I am the one that brought you out of the house of silence, where you were beating your chest and yelling in the forest. Into a house flowing with drums and guitars.
2 Thou shalt not grift.
3 Thou shalt not use chatbots to write on our forum. People use forums to talk to people. If they wanted to chat with a bot they would go to their favorite chatbot website.