# Using Nsight Systems to profile GPU workload

**URL:** <https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59>\
**Category:** NVIDIA CUDA\
**Created:** [January 25, 2021, 11:09am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59 "2021-01-25T11:09:04Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![ptrblck](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/ptrblck/32/27_2.png) [@ptrblck](https://dev-discuss.pytorch.org/u/ptrblck)\
**Post date:** [January 25, 2021, 11:09am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/1 "2021-01-25T11:09:04Z")

</div>

This topic describes a common workflow to profile workloads on the GPU using Nsight Systems.

As an example, let’s profile the `forward`, `backward`, and `optimizer.step()` methods using the `resnet18` model from `torchvision`.

To annotate each part of the training we will use `nvtx` ranges via the `torch.cuda.nvtx.range_push/.range_pop` operations. These ranges work as a stack and can be nested.  
Also, we are usually not interested in the first iteration, which might add overhead to the overall training due to memory allocations, cudnn benchmarking etc., thus we start the profiling after a few iterations via `torch.cuda.cudart().cudaProfilerStart()` and stop it at the end via `.cudaProfilerStop()`.

A complete code snippet can be seen here:

```python
import torch
import torch.nn as nn
import torchvision.models as models

# setup
device = 'cuda:0'
model = models.resnet18().to(device)
data = torch.randn(64, 3, 224, 224, device=device)
target = torch.randint(0, 1000, (64,), device=device)
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

nb_iters = 20
warmup_iters = 10
for i in range(nb_iters):
    optimizer.zero_grad()

    # start profiling after 10 warmup iterations
    if i == warmup_iters: torch.cuda.cudart().cudaProfilerStart()

    # push range for current iteration
    if i >= warmup_iters: torch.cuda.nvtx.range_push("iteration{}".format(i))

    # push range for forward
    if i >= warmup_iters: torch.cuda.nvtx.range_push("forward")
    output = model(data)
    if i >= warmup_iters: torch.cuda.nvtx.range_pop()

    loss = criterion(output, target)

    if i >= warmup_iters: torch.cuda.nvtx.range_push("backward")
    loss.backward()
    if i >= warmup_iters: torch.cuda.nvtx.range_pop()

    if i >= warmup_iters: torch.cuda.nvtx.range_push("opt.step()")
    optimizer.step()
    if i >= warmup_iters: torch.cuda.nvtx.range_pop()

    # pop iteration range
    if i >= warmup_iters: torch.cuda.nvtx.range_pop()

torch.cuda.cudart().cudaProfilerStop()

```

To create the profile I’m using Nsight System 2020.4.3.7 via the CLI.  
The CLI options for `nsys profile` can be found [here](https://docs.nvidia.com/nsight-systems/UserGuide/index.html#cli-profile-command-switch-options) and my “standard” command as well as the one used to create the profile for this example is:

```auto
nsys profile -w true -t cuda,nvtx,osrt,cudnn,cublas -s cpu --capture-range=cudaProfilerApi --stop-on-range-end=true --cudabacktrace=true -x true -o my_profile python main.py

```

(Thanks to Michael Carilli to create this cmd a while ago 😉 )

The arguments can be found in the linked CLI docs. A few interesting arguments are:

- `-t cuda,nvtx,osrt,cudnn,cublas`: selects the APIs to be traced
- `--capture-range=cudaProfilerApi` and `--stop-on-range-end=true`: profiling will start only when `cudaProfilerStart` API is invoked / Stop profiling when the capture range ends.
- `--cudabacktrace=true`: When tracing CUDA APIs, enable the collection of a backtrace when a CUDA API is invoked. (allows you to hover over a call and get the backtrace) You can also specify thresholds in `ns` which defines a threshold which the kernel must execute before backtraces are collected.

Note, if you are used to `nvprof`, try to copy/paste your `nvprof` cmd via:

```auto
nsys nvprof [options]

```

and Nsight Systems would try to translate the legacy `nvprof` command.

For this example the profile would look like this on a TitanV:

 ![img01](https://canada1.discourse-cdn.com/flex036/uploads/pytorch1/original/1X/b3d62d80fee5646484c48abca6207314034abab8.png)

You can see the execution in different Python threads as well as different APIs executing kernels on the device.  
We can zoom into a specific iteration and check the backtrace option, which is often helpful to isolate specific regressions and see which functions were invoked.

 ![img02](https://canada1.discourse-cdn.com/flex036/uploads/pytorch1/original/1X/0f87e630c4390a32402de95a9883abe5c12d3a7e.png)

Feel free to add more useful arguments or any tips and tricks using Nsight Systems.

---

<div class="post-metadata">

**Author:** ![mcarilli](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/mcarilli/32/14_2.png) [@mcarilli](https://dev-discuss.pytorch.org/u/mcarilli)\
**Post date:** [January 25, 2021, 7:39pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/2 "2021-01-25T19:39:47Z")

</div>

I have a gist with my preferred nsys commands for different scenarios and an explanation of each option.

> <https://gist.github.com/mcarilli/376821aa1a7182dfcf59928a7cde3223>

There’s a gotcha to be aware of with Piotr’s command line

```auto
nsys profile -w true -t cuda,nvtx,osrt,cudnn,cublas -s cpu --capture-range=cudaProfilerApi --stop-on-range-end=true --cudabacktrace=true -x true -o my_profile python main.py

```

CPU sampling (`-s cpu`) is great for getting backtraces that shows where particular timeline calls originate in the code, but also inflates CPU overhead (sometimes dramatically, 2X or more). So with `-s cpu`, you shouldn’t expect a realistic view of CPU whitespace.

To get a better idea of bottlenecks, you should first create a profile without CPU sampling (`-s none`), eg

```auto
nsys profile -w true -t cuda,nvtx,osrt,cudnn,cublas -s none -o nsight_report -f true -x true python script.py args...

```

Manual `torch.cuda.nvtx.range_push/pop` calls in your script are very helpful to orient yourself and immediately see where your code spends an unexpected amount of time (forward? backward? optimizer? between iterations, which usually means dataloader?).

---

<div class="post-metadata">

**Author:** ![zaccharieramzi](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/zaccharieramzi/32/176_2.png) [@zaccharieramzi](https://dev-discuss.pytorch.org/u/zaccharieramzi)\
**Post date:** [May 21, 2021, 4:56pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/3 "2021-05-21T16:56:22Z")

</div>

Hi,

Is this supposed to work with multi-GPU scripts? I am using DataParallel in my case.

Right now I have the following feedback:

```auto
The application terminated before the collection started. No report was generated.
	Collection canceled.

```

Thanks in advance

---

<div class="post-metadata">

**Author:** ![ekonyagin](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/ekonyagin/32/356_2.png) [@ekonyagin](https://dev-discuss.pytorch.org/u/ekonyagin)\
**Post date:** [December 17, 2021, 6:05pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/4 "2021-12-17T18:05:53Z")

</div>

Hi,

is it necessary to conduct device synchronize if am interested in CUDA Kernel time for each op of the layer?

---

<div class="post-metadata">

**Author:** ![pyotr777](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/pyotr777/32/459_2.png) [@pyotr777](https://dev-discuss.pytorch.org/u/pyotr777)\
**Post date:** [March 7, 2022, 9:49am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/5 "2022-03-07T09:49:02Z")

</div>

In newer PyTorch/nsys versions I don’t see cuDNN info anymore.  
Can anyone help with that?

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/pytorch1/original/1X/d2b40fefc23729e32cdd63bda30ffba12c9bf3d4.png)

I also posted on NVIDIA forums here:

> **[No cuDNN info in nsys traces](https://forums.developer.nvidia.com/t/no-cudnn-info-in-nsys-traces/204630)**
>
> In recent versions of nsys, I see no cuDNN info. Instead, I see errors cuDNN profiling might have not been started correctly. I am using the following CLI options: $ sudo nsys profile -t cudnn,cuda,nvtx,osrt -o report\_%h\_%p python3 ... In earlier...

---

<div class="post-metadata">

**Author:** ![pyotr777](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/pyotr777/32/459_2.png) [@pyotr777](https://dev-discuss.pytorch.org/u/pyotr777)\
**Post date:** [May 16, 2022, 6:56am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/6 "2022-05-16T06:56:10Z")

</div>

If I build PyTorch from sources, then there is cuDNN info in nsys traces.

---

<div class="post-metadata">

**Author:** ![fxmarty](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/fxmarty/32/828_2.png) [@fxmarty](https://dev-discuss.pytorch.org/u/fxmarty)\
**Post date:** [December 26, 2022, 10:47pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/7 "2022-12-26T22:47:29Z")

</div>

> We can zoom into a specific iteration and check the backtrace option, which is often helpful to isolate specific regressions and see which functions were invoked.

I can’t quite manage to get the backtrace feature to work, most of the kernels have no call stack when hovering over them. Did anyone encounter the issue? Posted on nvidia forums as well to see what’s wrong [Call stack is visible/captured only for some CUDA kernels (broken backtraces) - Profiling Linux Targets - NVIDIA Developer Forums](https://forums.developer.nvidia.com/t/call-stack-is-visible-captured-only-for-some-cuda-kernels/238005)

I’m wondering if I should not build pytorch from source with `-fno-omit-frame-pointer` for this to work.

---

<div class="post-metadata">

**Author:** ![shubbey](https://avatars.discourse-cdn.com/v4/letter/s/e495f1/32.png) [@shubbey](https://dev-discuss.pytorch.org/u/shubbey)\
**Post date:** [February 24, 2023, 8:40pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/8 "2023-02-24T20:40:30Z")

</div>

I have noticed that when profiling my networks with nsys, the cpu is always running 100% during loss.backward(). The graph looks similar to the one in the first picture here. I was hoping someone could explain what is happening here, because I am trying to find bottlenecks in my routine as I don’t seem to be gettng expected performance boosts when increasing the batch size, using AMP, and so forth. It also doesn’t seem to matter whether I preload my data onto the gpu or use workers to transfer it from the cpu at runtime. In both cases, I only have appreciable cpu activity during loss.backward(). I was under the impression that if I preloaded all of my data and only retrieved shuffled batches at training time via, eg. torch.gather(), I would not be using the cpu at all. Any help? Thanks!

---

<div class="post-metadata">

**Author:** ![jendrikjoe](https://avatars.discourse-cdn.com/v4/letter/j/a698b9/32.png) [@jendrikjoe](https://dev-discuss.pytorch.org/u/jendrikjoe)\
**Post date:** [March 29, 2023, 12:29pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/9 "2023-03-29T12:29:55Z")

</div>

Hey folks, thanks for starting this thread. It proofed very helpful in trying to profile my application. I am currently following the PyTorch lightning guide: [Find bottlenecks in your code (intermediate) — PyTorch Lightning 2.0.0 documentation](https://lightning.ai/docs/pytorch/stable/tuning/profiler_intermediate.html#visualize-profiled-operations) and use `nsys profile -w true -t cuda,nvtx,osrt,cudnn,cublas -s none --capture-range-end stop --capture-range=cudaProfilerApi --cudabacktrace=true -x true poetry run python main_graph.py` as the command to collect the emitted information. However, I am getting a 300MB file just when doing one step of training and a 1.7GB file doing thirty steps. Can anyone give me a hint if this is due to a huge misconfiguration on my side, or owed to the fact of using PyTorch lighting and PyTorch geometric?  
Any input would be greatly appreciated 🤗

---

<div class="post-metadata">

**Author:** ![LukeLIN-web](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/lukelin-web/32/1579_2.png) [@LukeLIN-web](https://dev-discuss.pytorch.org/u/LukeLIN-web)\
**Post date:** [October 30, 2023, 12:59pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/10 "2023-10-30T12:59:03Z")

</div>

In nsight 2023.3.1 , one option is outdated.

```auto
nsys profile -w true -t cuda,nvtx,cudnn,cublas --capture-range=cudaProfilerApi --stop-on-range-end=true -x true -o 512 python ladies_e2e.py
unrecognised option '--stop-on-range-end=true'

usage: nsys profile [<args>] [application] [<application args>]
Try 'nsys profile --help' for more information.

```

---

<div class="post-metadata">

**Author:** ![0324wy](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/0324wy/32/1980_2.png) [@0324wy](https://dev-discuss.pytorch.org/u/0324wy)\
**Post date:** [May 24, 2024, 3:26am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/11 "2024-05-24T03:26:39Z")

</div>

Yes, I got the same error.

---

<div class="post-metadata">

**Author:** ![yasin](https://avatars.discourse-cdn.com/v4/letter/y/9e8a1a/32.png) [@yasin](https://dev-discuss.pytorch.org/u/yasin)\
**Post date:** [August 22, 2024, 3:23pm UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/12 "2024-08-22T15:23:44Z")

</div>

You can use `--capture-range-end=stop` instead.

---

<div class="post-metadata">

**Author:** ![crcrpar](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/crcrpar/32/152_2.png) [@crcrpar](https://dev-discuss.pytorch.org/u/crcrpar)\
**Post date:** [April 30, 2025, 6:10am UTC](https://dev-discuss.pytorch.org/t/using-nsight-systems-to-profile-gpu-workload/59/13 "2025-04-30T06:10:57Z")

</div>

In recent nsight systems, probably from 2025.1 as per [Release Notes — nsight-systems 2025.1 documentation](https://docs.nvidia.com/nsight-systems/2025.1/ReleaseNotes/index.html), there’s an option specific for pytorch, `--pytorch=autograd-nvtx` and so on – [User Guide — nsight-systems 2025.2 documentation](https://docs.nvidia.com/nsight-systems/UserGuide/index.html#pytorch-profiling).
