DPC, Cubase, Windows

I have same behavior with ntoskrnl.exe. It ramps to 600-700 every 10-15 min.
Win 11 pro 23h2
Asus ROG Strix Z790-e
13900ks

Seems there is something going on with this thing.
Maybe @Psychlist1972 could help with this NTOSKRNL.EXE weird spike and check this?

I have otherwise absolutely perfect Latency Mon results and only annoying thing is that NTOSKRNL.EXE spike repeating every 10-15 min interval. @Psychlist1972

That file is the OS kernel process, I have not yet determined what is going on.

Also, it’s quite literally every 15 minutes +/- 30 seconds or so.

So far, I have not been able to detect an interrupt starting it, and as I mentioned, the other things going on appear nondeterministic. I have yet to isolate the common thread causing it.

With the interval being NEARLY constant, it appears to be a software scheduled task. It happens on another machine with Windows 10, just shorter duration and may possibly be related to the number of cores.

Given the slight variability in the interval, it could potentially be a lower priority task, but that is simply a guess at this point.

I’m mulling over how to further capture what’s going on.

If you’re using Process Lasso, I will include a couple of other changes regarding getting audio tasks off of CPU 0/1 as well as increasing the priority and I/O priority of Cubase itself to Higher (and also off of CPU 0/1). Stay tuned! I’m not convinced this is needed, but those are changes that got my Cubase useable while I discovered more information.

Well firstly do not do a “Diagnostic restart” using ms config.
You can’t sign in and have to use system restore to get back to Windows through recovery. A warning would have been helpful, Latencymon won’t run in Safe mode either.
I agree with your timings, it is around 15 minutes but not exactly 15 minutes.
At present my suspicion is Intel or Windows.
In the posted video he talks about moving the operating system off core 0, and mentions other tweaks without saying what they are.
Further investigation needed !

I would characterize it as 15 minutes with some variation (as in NOT exact). It’s pretty darn close and often, right on the millisecond. I may still have a one-hour trace that shows it clearly. On one trace I clearly saw two right on 15 minutes, with the third a few seconds off followed by the fourth almost right on. For practical purposes, it’s 15 minutes NOT 10-15.

That’s not to say there might be something else going on with your scenario causing something else, but two machines (one a dual-boot Windows 10/11) show it reliably at 15 minutes +/- a small error.

There are tons of videos suggesting all sorts of things. Some I can speculate why might make improvements, but I’m not confident they address the root cause. I am NOT interested in throwing changes against the wall to see what sticks. I want to evaluate against the likely underlying cause.

At some point I will evaluate the changes you mentioned and test with the ones that make sense against the data results. That may not happen today!

I am trying to be quite purposeful in understanding WHY changes are being made and the underlying theory of the problem. I don’t know if we’ll have enough visibility of the issue without some help. SOMEBODY has to know what’s being scheduled every 15 minutes! My guess is some OS housekeeping that should probably be broken into smaller real-time segments.

I’m also trying not to put the cart in front of the horse, addressing SYMPTOMS rather than the cause.

Your confirmation of the data and some things you’ve tried is helpful. I want to take care not to get detoured from the root cause.

Understand your position, but would suggest that the fact it happens more than once is probably due to whatever is causing it not functioning properly and so it tries again.
I ran Latencymon for 40 minutes and there were 8 instances of it, I only visibly saw one.
Fingers crossed somebody with more knowledge helps.
I suspect there are many that have the problem but are unaware of it.
Also if you are using the Cubase power plan you need to make the power plan changes with Cubase open.
Best of luck and happy Christmas.
Now I’m going back to throw some more stuff at the wall :grinning:.

Maybe, I don’t think that’s what’s happening but cannot prove yet, so…

I’ve run this testing a LOT, and the results vary a little in duration, but zooming in shows the two events are consistent, only the duration changes, and of course whatever happens to be running at the time, which can easily be misconstrued as “random”. Could it be a failure and retry? Yes, it COULD be, I just don’t think that’s what we’re seeing.

My hope here is that zooming in on the problem and providing some data will ultimately filter to some folks internally and trigger some curiosity, and perhaps even some knowledge of how to isolate the cause.

I appreciate the info you posted, just have not spent the time to verify to my satisfaction. I’m also working with my internal network and Comcast speed issues that are AT LEAST as frustrating associated with the Netgear CM2050V modem. As usual, I have to PROVE it’s not my issue, even though they have knowingly identified firmware issues AND pushed out updates and downgrades.

I may get time to summarize in one place what I’ve done so far and compare with your suggestions as well (unrelated to the speed issues).

I would suggest you try not to throw the hardware at the wall, as tempting as it becomes! :smiley:

The spike may well be caused by the background process scheduler for UWP apps, which wakes up once every 15 minutes. Apps like Mail, Phone Link, XBox, Disney+, etc. are UWP apps which are dependent on this scheduler if they have an associated background process on a timer.
The spike happens on our machines here as well but the DPC latencies are not significant enough to have any effect on Cubase / Nuendo performance so, I ignore it.
I’m not saying this is the answer - it is one possibility. I haven’t spent any time investigating it.

I think this makes sense and also makes this one potential candidate. The variations in the 15 minute period could potentially be explained by the amount of mail queued and the scheduling of the next 15 minute AFTER processing the queue(s).

This is exactly the type of “soft” scheduling that might account for variations.

ANY candidate that is known to occur on a 15-minute interval is potentially a suspect until we have more data. In the case of email, I would think that some adjustments to processor affinity and/or priorities could potentially solve the issue(s). Still too early to say, but potentially useful info.

I have just changed all the processes in ProcessLasso to idle under Priority, some don’t allow it though, no difference, it still happens.

Not sure if this workaround I stumbled across will be relevant to your situation but have you tried disabling “hardware accelerated GPU scheduling” on your system? I was seeing unacceptable DPC latency spikes in Latencymon with ntoskernl.exe & my NVIDIA card driver before I disabled this switch.

Modern AMD cards appear not to have this setting.
From what I understand the Radeon cards have GPU Scheduling via Hardware so it is not possible to disable the windows gpu scheduling because it is no longer used.

I have tried moving all running processes away from core 0 and 1, no effect on the relevant spike.
I tried running Latency Mon behind Cinebench 23, `and even though LM was slightly unresponsive spike was still there.
So this may mean that it’s not a process that needs the PC to be running in idle.
An article I found may help in tracking it down if you are using the Windows Assessment & deployment kit.
https://mahdytech.com/livesite-thousands-spikes/
To me it means very little but I’m hoping that in it there may be clues as to how to track down a 15 minute latency spike.

Thank you for sharing this, does it mean that it has the opposite effect of what it supposed to do?

Thanks, but again, I’m looking for Steinberg, Yamaha, and other companies in the music space with deep pockets and plenty of time to band together and develop the tools & methods to fix this stuff.

Even a process viewer from within Cubase, something that shows exactly what plugin is doing what to the CPU would be hugely beneficial for us.

Their priorities are elsewhere, so my purpose is to make a lot of noise and emphasize that the status quo is really unacceptable, that it’s gone on long enough.

That is my point: very, very tired of spending so much time trying different bullshit that I shouldn’t have to try in the first place. This is an engineering problem, first and foremost, and I’m not being paid to be an engineer for Steinberg.

I’m fed up doing free beta testing in production and free troubleshooting for these incompetents.

AMEN Dr!!!

The 15 minute latency spike this thread relates to is happening on windows without any music software running.
So not a Cubase or Yamaha problem.
But under heavy load it will effect latency in Cubase.
But what you say is true, in general there are tools that help but they often lead you nowhere even if they give you hints, such as ntoskrnl.exe and it happens every 15 minutes.
Still no success at finding out what’s causing it, unfortunately.

I’ve isolated the every 15 minutes spike using WPA. The problems seems to be related the function ExpGetPoolTagInfoTarget. I couldn’t get much info about it other than the fact that looks like the kernel is collecting some driver’s data from memory.

A pool tag is a four-byte character that is associated with a dynamically allocated chunk of pool memory. The tag is specified by a driver when it allocates the memory. So how can you figure out which tag belongs to which driver? There is a file ( Pooltag.txt ) that lists the pool tags used for pool allocations by kernel-mode components and drivers supplied with Windows.

I can’t post links here, but you can find pooltag.txt online without issues. It also comes with Windows Debugging Tools for Windows.

Anyway, here is the screenshot from WPA.

ReactOS (a free, opensource reimplementation of windows) implements that function as this but not much more info is given…

VOID
NTAPI
ExpGetPoolTagInfoTarget(IN PKDPC Dpc,
                        IN PVOID DeferredContext,
                        IN PVOID SystemArgument1,
                        IN PVOID SystemArgument2)
{
    PPOOL_DPC_CONTEXT Context = DeferredContext;
    UNREFERENCED_PARAMETER(Dpc);
    ASSERT(KeGetCurrentIrql() == DISPATCH_LEVEL);
 
    //
    // Make sure we win the race, and if we did, copy the data atomically
    //
    if (KeSignalCallDpcSynchronize(SystemArgument2))
    {
        RtlCopyMemory(Context->PoolTrackTable,
                      PoolTrackTable,
                      Context->PoolTrackTableSize * sizeof(POOL_TRACKER_TABLE));
 
        //
        // This is here because ReactOS does not yet support expansion
        //
        ASSERT(Context->PoolTrackTableSizeExpansion == 0);
    }
 
    //
    // Regardless of whether we won or not, we must now synchronize and then
    // decrement the barrier since this is one more processor that has completed
    // the callback.
    //
    KeSignalCallDpcSynchronize(SystemArgument2);
    KeSignalCallDpcDone(SystemArgument1);
}

I think we could get more info, such as arguments used in the invocation, putting a breakpoint with WinDgb or something like that but I’ve never used it before and should invest some more time learning.

Anyone else could help?

After some reverse engineering I traced that every 15 minutes or so, it is svchost.exe that creates a new thread to log all pool tag info (it ends up calling ExpGetPoolTagInfoTarget via DPC). As svhost.exe is used to host and manage Windows services I bet there’s one service that we could potentially stop and solve the issue. Let see if we can find it…