# PyTorch 2.x Inference Recommendations

**URL:** <https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506>\
**Category:** deployment\
**Created:** [September 30, 2024, 5:23pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506 "2024-09-30T17:23:52Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![agunapal](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/agunapal/32/810_2.png) [@agunapal](https://dev-discuss.pytorch.org/u/agunapal)\
**Post date:** [September 30, 2024, 5:23pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/1 "2024-09-30T17:23:52Z")

</div>

## PyTorch 2.x Inference Recommendations

PyTorch 2.x introduces a range of new technologies for model inference and it can be overwhelming to figure out which technology is most appropriate for your particular use case. This guide aims to provide clarity and guidance on the various options available.

### `torch.compile`

`torch.compile` speeds up PyTorch code by JIT compiling PyTorch code into optimized kernels with a simple API . It optimizes the given model using TorchDynamo and creates an optimized graph , which is then lowered into the hardware using the backend specified in the API. The default backend in torch.compile is [TorchInductor](https://dev-discuss.pytorch.org/t/torchinductor-a-pytorch-native-compiler-with-define-by-run-ir-and-symbolic-shapes/747). For more information, please refer to this [tutorial](https://pytorch.org/tutorials/intermediate/torch_compile_tutorial.html)

### AOTInductor (CPP)

AOTInductor(AOTI) is a specialized version of TorchInductor that takes a PyTorch exported program , optimizes it and produces a shared library artifact that can be used in a non-Python deployment environment. We use `torch.export` to capture the model into a computational graph and then use AOTCompile to generate the shared library which can be loaded and executed in a CPP deployment environment. More details can be found in this [tutorial](https://pytorch.org/docs/stable/torch.compiler_aot_inductor.html)

### AOTInductor (Python)

When the shared library generated by AOTI needs to be executed in Python runtime, AOTI provides a python API to do so. `torch._export.aot_load` is used to load and execute the shared library generated by AOTCompile. For further details, please consult this [tutorial](https://pytorch.org/tutorials/recipes/torch_export_aoti_python.html)

### ExecuTorch

ExecuTorch is an end-to-end PyTorch platform that provides the infrastructure to run PyTorch programs on edge. The devices can range from AR/VR wearables to Android/iOS mobile devices. It relies heavily on `torch.compile` and `torch.export`. ExecuTorch provides exhaustive [documentation](https://pytorch.org/executorch/stable/index.html) on the end-to-end flow for various hardware platforms.

In the below table, you can find PyTorch’s recommendations on which technology to use depending on the platform where the inference is being done and the use case where this deployment is being used. Please note that some of the export related APIs mentioned above could change.

### Inference Recommendations

| No. | Platform | Use Case | Recommendation |
| --- | --- | --- | --- |
| 1 | Mobile (iOS and Android) | Highly Optimized Inference | ExecuTorch |
| 2 | Embedded and other Edge Devices | Highly Optimized Inference | Executorch |
| 3 | Server (x86, CUDA, aarch64 ) + Python deployment | High Throughput, highly optimized inference with no graph break | AOTI (Python) |
| 4 | Server (x86, CUDA, aarch64) + Python deployment | High Throughput, highly optimized inference with graph break | `torch.compile` |
| 5 | Server (when AOTI is not supported) + Python deployment | High Throughput, highly optimized inference | `torch.compile` |
| 6 | Server (x86, CUDA, aarch64) with non-Python deployment | High Throughput, highly optimized inference with no graph break | AOTI (CPP) |
| 7 | Server (x86, CUDA, aarch64) + Python deployment | When non inductor backend has the best latency | `torch.compile` |
| 8 | Mac (M1/M2/M3) | Local development and inference | eager |

---

<div class="post-metadata">

**Author:** ![spiegelball](https://avatars.discourse-cdn.com/v4/letter/s/5f9b8f/32.png) [@spiegelball](https://dev-discuss.pytorch.org/u/spiegelball)\
**Post date:** [October 15, 2024, 4:14pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/2 "2024-10-15T16:14:18Z")

</div>

What is the recommended solution for non-Python mac systems? Is torchscript still the way to go?

---

<div class="post-metadata">

**Author:** ![smth](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/smth/32/3_2.png) [@smth](https://dev-discuss.pytorch.org/u/smth)\
**Post date:** [October 15, 2024, 5:43pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/3 "2024-10-15T17:43:36Z")

</div>

non-Python mac systems would be AOTI too.

TorchScript is deprecated and not the way to go.

---

<div class="post-metadata">

**Author:** ![spiegelball](https://avatars.discourse-cdn.com/v4/letter/s/5f9b8f/32.png) [@spiegelball](https://dev-discuss.pytorch.org/u/spiegelball)\
**Post date:** [October 16, 2024, 9:58am UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/4 "2024-10-16T09:58:04Z")

</div>

Do you have any sources on AOTI inference on mac? I only found this: [AOT Inductor and macOS · Issue #119803 · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/issues/119803)

And regarding TorchScript, the latest info I saw is coming from this thread:

> [@What's the difference between torch.export / torchserve / executorch / aotinductor?](https://dev-discuss.pytorch.org/t/whats-the-difference-between-torch-export-torchserve-executorch-aotinductor/1642/15):
>
> We will not deprecate TorchScript without a suitable (and technically superior) replacement. I think the key missing piece, which we are developing but have not yet released, is a generic interpreted runtime that uses libtorch to execute the graph in a target-independent way, optionally calling out to compiled artifacts for acceleration. So the proposed TorchScript replacement flow would be: On the frontend: torch.export → compile subgraphs/whole graph with inductor → packaged model (graph, …

Is it deprecated now?

---

<div class="post-metadata">

**Author:** ![agunapal](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/agunapal/32/810_2.png) [@agunapal](https://dev-discuss.pytorch.org/u/agunapal)\
**Post date:** [October 16, 2024, 3:42pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/5 "2024-10-16T15:42:27Z")

</div>

The same tutorial mentioned in the guide works on mac (cpu) [torch.export AOTInductor Tutorial for Python runtime (Beta) — PyTorch Tutorials 2.4.0+cu121 documentation](https://pytorch.org/tutorials/recipes/torch_export_aoti_python.html)

---

<div class="post-metadata">

**Author:** ![spiegelball](https://avatars.discourse-cdn.com/v4/letter/s/5f9b8f/32.png) [@spiegelball](https://dev-discuss.pytorch.org/u/spiegelball)\
**Post date:** [October 16, 2024, 4:49pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/6 "2024-10-16T16:49:09Z")

</div>

Thanks for the link, but this is for python-based inference, whereas the systems in question are non-python.

---

<div class="post-metadata">

**Author:** ![agunapal](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/agunapal/32/810_2.png) [@agunapal](https://dev-discuss.pytorch.org/u/agunapal)\
**Post date:** [October 17, 2024, 6:37pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/7 "2024-10-17T18:37:47Z")

</div>

You can take a look at how torchchat does this on mac [GitHub - pytorch/torchchat: Run PyTorch LLMs locally on servers, desktop and mobile](https://github.com/pytorch/torchchat?tab=readme-ov-file#run-using-our-c-runner)

---

<div class="post-metadata">

**Author:** ![bhack](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/bhack/32/2339_2.png) [@bhack](https://dev-discuss.pytorch.org/u/bhack)\
**Post date:** [October 22, 2024, 2:15pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/8 "2024-10-22T14:15:56Z")

</div>

I think that it will be also useful for the AOT case to have a documentation clarification over this old but good thread:

> <https://github.com/pytorch/pytorch/issues/115965>
>
> \### 🚀 The feature, motivation and pitch
> 
> AOT inductor looks like the upcoming me…ans to do inference from native code that was trained in pytorch, and the replacement for torchcript export to native code. It's clear this interface is in \[prototype status\](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler\_aot\_inductor.rst), but based on what is present right now, it's problematic for many users. 
> 
> torch.\_export.aot\_compile, as currently defined, produces a .so and presumably invokes nvcc for gpu models, and likely a host compiler for cpu models. This is pretty problematic for integration into many native build tools, as the export process takes over building of the inference library. Cross compilation is impossible, as is passing flags to the build tools. 
> 
> An interface that would potentially be much friendlier would yield source code that could be fed into an existing build system rather than directly providing a library. This way pytorch wouldn't have to manage build tools in any capacity. There is likely a lot of complexity here, because code generation likely wants to hardcode many platform specific details, e.g. gpu type, cpu instruction sets. Being able to specify the platform constraints and capabilities on the aoti\_compile interface would likely be wise, rather than attempting to automatically infer it from the local machine. 
> 
> 
> \### Alternatives
> 
> \_No response\_
> 
> \### Additional context
> 
> \_No response\_
> 
> cc @ezyang @anijain2305 @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @peterbell10 @ipiszy @yf225 @chenyang78 @kadeng @muchulee8 @ColinPeppler @amjames @desertfire @msaroufim @wconstab @bdhirsh @zou3519 @aakhundov

Also partially related to this from the export tutorial:

> As `torch.export` is only a graph capturing mechanism, calling the artifact produced by `torch.export` eagerly will be equivalent to running the eager module. To optimize the execution of the Exported Program, we can pass this exported artifact to backends such as Inductor through `torch.compile`, [AOTInductor](https://pytorch.org/docs/main/torch.compiler_aot_inductor.html), or [TensorRT](https://pytorch.org/TensorRT/dynamo/dynamo_export.html).

It is quite confusing what kind of benefit we could have using `torch.export` and the different `torch.compile` options especially in the case where we are exporting on an host with a different hardware from the target inference host.

---

<div class="post-metadata">

**Author:** ![agunapal](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/agunapal/32/810_2.png) [@agunapal](https://dev-discuss.pytorch.org/u/agunapal)\
**Post date:** [October 22, 2024, 8:45pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/9 "2024-10-22T20:45:25Z")

</div>

Regarding your second point, I believe the user experience would be to create the shared library on the same hardware where it would be deployed. I believe this was the same in TorchScript?

And in typical deployments, a given model would be deployed on one kind of hardware. So, this means there is only 1 shared library per model, which shouldn’t make the model ops complicated?

---

<div class="post-metadata">

**Author:** ![bhack](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/bhack/32/2339_2.png) [@bhack](https://dev-discuss.pytorch.org/u/bhack)\
**Post date:** [October 22, 2024, 9:17pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/10 "2024-10-22T21:17:22Z")

</div>

> And in typical deployments, a given model would be deployed on one kind of hardware.

I believe this point relates more closely to the first issue and the referenced ticket/thread.

Regarding the second point, users may still feel a bit confused about the differences between AOTI performance/optimization and regular `torch.compile`.

For instance, consider a scenario where we have a discrete number of inputs with dynamic sizes that aren’t limited to the batch dimension. This situation often presents a trade-off between padding and targeting a discrete number of input dimensions, which is particularly common in vision tasks with varying input resolutions during inference.

With the AOTI approach during export, you can only specify a range of values without the option to define a finite set of actual sparse inputs. You can see more on this in [GitHub Issue #136119](https://github.com/pytorch/pytorch/issues/136119).

In contrast, using the “classical” `torch.compile` allows you to generate specific code tailored to the exact types of inputs you’ll be feeding into the model.

Some users might prefer to optimize the code for a finite set of dimensions while still leveraging the `torch.compile` cache or remote cache.

However, it appears to be challenging for users to grasp the performance benefits of the two solutions, especially when compiling/exporting on the same inference hardware target.

---

<div class="post-metadata">

**Author:** ![spiegelball](https://avatars.discourse-cdn.com/v4/letter/s/5f9b8f/32.png) [@spiegelball](https://dev-discuss.pytorch.org/u/spiegelball)\
**Post date:** [October 24, 2024, 10:59am UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/11 "2024-10-24T10:59:10Z")

</div>

No, torchscript would create one artifact which could run on any other hardware as well.

---

<div class="post-metadata">

**Author:** ![ipickering](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/ipickering/32/2355_2.png) [@ipickering](https://dev-discuss.pytorch.org/u/ipickering)\
**Post date:** [November 3, 2024, 2:13pm UTC](https://dev-discuss.pytorch.org/t/pytorch-2-x-inference-recommendations/2506/12 "2024-11-03T14:13:13Z")

</div>

Is there current or planned support for autograd on AOTI and `torch.export`? By the way, excellent explanation, it is a bit overwhelming, but I’m excited for the new features =)
