# compiler

**URL:** https://dev-discuss.pytorch.org/c/compiler/5.md

[Latest](https://dev-discuss.pytorch.org/latest.md) · [Categories](https://dev-discuss.pytorch.org/categories.md)

---

## [About the compiler category](https://dev-discuss.pytorch.org/t/about-the-compiler-category/14)

<div class="topic-metadata">

**Author:** [@smth](https://dev-discuss.pytorch.org/u/smth)\
**Replies:** 0\
**Last updated:** [January 22, 2021, 7:19pm UTC](https://dev-discuss.pytorch.org/t/about-the-compiler-category/14 "2021-01-22T19:19:40Z")

</div>

---

## [\[RFC-0036\] Zero-GC 64-Byte Cache-Aligned Flat Arena for Speculative Decoding Verification](https://dev-discuss.pytorch.org/t/rfc-0036-zero-gc-64-byte-cache-aligned-flat-arena-for-speculative-decoding-verification/3441)

<div class="topic-metadata">

**Author:** [@markbgilbert](https://dev-discuss.pytorch.org/u/markbgilbert)\
**Replies:** 0\
**Last updated:** [September 21, 2026, 4:26pm UTC](https://dev-discuss.pytorch.org/t/rfc-0036-zero-gc-64-byte-cache-aligned-flat-arena-for-speculative-decoding-verification/3441 "2026-09-21T16:26:22Z")

</div>

We have submitted RFC-0036: Zero-GC 64-Byte Cache-Aligned Flat Arena Runtime to address host-side CPU overhead and allocator contention in high-throughput LLM speculative decoding and batch token ingestion. Empirical Pr…

---

## [Review request: custom-kernel attribution benchmarks vs torch.compile for decode workloads](https://dev-discuss.pytorch.org/t/review-request-custom-kernel-attribution-benchmarks-vs-torch-compile-for-decode-workloads/3411)

<div class="topic-metadata">

**Author:** [@MilkClouds](https://dev-discuss.pytorch.org/u/MilkClouds)\
**Replies:** 0\
**Last updated:** [July 1, 2026, 5:54am UTC](https://dev-discuss.pytorch.org/t/review-request-custom-kernel-attribution-benchmarks-vs-torch-compile-for-decode-workloads/3411 "2026-07-01T05:54:04Z")

</div>

Hi PyTorch compiler folks, I put together a small benchmark repo on custom-kernel attribution for decode/generation workloads: The narrow claim is not that custom kernels are useless. The claim is that in these publi…

---

## [RFC: Polyhedral Optimization Pass for PyTorch Inductor](https://dev-discuss.pytorch.org/t/rfc-polyhedral-optimization-pass-for-pytorch-inductor/3341)

<div class="topic-metadata">

**Author:** [@morrison-turnansky](https://dev-discuss.pytorch.org/u/morrison-turnansky)\
**Replies:** 2\
**Last updated:** [June 9, 2026, 7:10pm UTC](https://dev-discuss.pytorch.org/t/rfc-polyhedral-optimization-pass-for-pytorch-inductor/3341 "2026-06-09T19:10:29Z")

</div>

RFC: Polyhedral Optimization Pass for PyTorch Inductor This RFC proposes adding an optional polyhedral optimization pass to PyTorch Inductor to enable fusion of operations that current fusion heuristics cannot handle. T…

---

## [\[RFC\] Improve Dynamic Shapes Support Across Aten Operators and Expand Test Coverage](https://dev-discuss.pytorch.org/t/rfc-improve-dynamic-shapes-support-across-aten-operators-and-expand-test-coverage/3392)

<div class="topic-metadata">

**Author:** [@kurator14](https://dev-discuss.pytorch.org/u/kurator14)\
**Replies:** 1\
**Last updated:** [June 8, 2026, 4:40pm UTC](https://dev-discuss.pytorch.org/t/rfc-improve-dynamic-shapes-support-across-aten-operators-and-expand-test-coverage/3392 "2026-06-08T16:40:13Z")

</div>

I’d like to propose a focused effort to improve dynamic shapes support across ATen operators and expand the corresponding test coverage. I would really appreciate any feedback on the scope, prioritization, or approach – …

---

## [Inductor Passes](https://dev-discuss.pytorch.org/t/inductor-passes/2742)

<div class="topic-metadata">

**Author:** [@jeromeku](https://dev-discuss.pytorch.org/u/jeromeku)\
**Replies:** 4\
**Last updated:** [June 3, 2026, 5:29pm UTC](https://dev-discuss.pytorch.org/t/inductor-passes/2742 "2026-06-03T17:29:02Z")

</div>

Is there any documentation on the sequence of fx passes that are run by Inductor? E.g., in the Writing Custom Backends Guide, dynamo passes a torch.fx.GraphModule to the backend. However, this graph is at a high level …

---

## [I keep getting this error while building PyTorch](https://dev-discuss.pytorch.org/t/i-keep-getting-this-error-while-building-pytorch/3351)

<div class="topic-metadata">

**Author:** [@YeonguChoe](https://dev-discuss.pytorch.org/u/YeonguChoe)\
**Replies:** 3\
**Last updated:** [May 8, 2026, 10:27pm UTC](https://dev-discuss.pytorch.org/t/i-keep-getting-this-error-while-building-pytorch/3351 "2026-05-08T22:27:53Z")

</div>

Hello I am planning to contribute to PyTorch project. I followed the installation process on the GitHub README. git submodule sync git submodule update --init --recursive python3 -m venv .venv source .venv/bin/activa…

---

## [Torch.compile as a toolkit / manipulating dynamo outputs](https://dev-discuss.pytorch.org/t/torch-compile-as-a-toolkit-manipulating-dynamo-outputs/3306)

<div class="topic-metadata">

**Author:** [@tom](https://dev-discuss.pytorch.org/u/tom)\
**Replies:** 4\
**Last updated:** [April 5, 2026, 12:49pm UTC](https://dev-discuss.pytorch.org/t/torch-compile-as-a-toolkit-manipulating-dynamo-outputs/3306 "2026-04-05T12:49:32Z")

</div>

Hi, so with great interest, I heard @ezyang ‘s comments about torch.compile as a toolkit at PTC. I’ve been trying to access/manipulate some of the things that dynamo produces. So there is a lot of material on what to d…

---

## [How is pattern matching in inductor/fx implemented?](https://dev-discuss.pytorch.org/t/how-is-pattern-matching-in-inductor-fx-implemented/1720)

<div class="topic-metadata">

**Author:** [@youkaichao](https://dev-discuss.pytorch.org/u/youkaichao)\
**Replies:** 9\
**Last updated:** [March 20, 2026, 4:20am UTC](https://dev-discuss.pytorch.org/t/how-is-pattern-matching-in-inductor-fx-implemented/1720 "2026-03-20T04:20:59Z")

</div>

Inductor and FX heavily use pattern matching to match and replace subgraph patterns for optimization. However, subgraph matching (finding a subgraph that is isomorphic to a given graph) is a well-known NP-hard problem. W…

---

## [Feature: Pattern Matcher Observability](https://dev-discuss.pytorch.org/t/feature-pattern-matcher-observability/3321)

<div class="topic-metadata">

**Author:** [@morrison-turnansky](https://dev-discuss.pytorch.org/u/morrison-turnansky)\
**Replies:** 1\
**Last updated:** [March 10, 2026, 3:44pm UTC](https://dev-discuss.pytorch.org/t/feature-pattern-matcher-observability/3321 "2026-03-10T15:44:03Z")

</div>

\# Pattern Matcher Observability ## Summary This RFC proposes adding opt-in debugging and observability features to the TorchInductor Patter Matcher. These features will help developers understand which patterns are app…

---

## [Helion Inductor Integration](https://dev-discuss.pytorch.org/t/helion-inductor-integration/3302)

<div class="topic-metadata">

**Author:** [@morrison-turnansky](https://dev-discuss.pytorch.org/u/morrison-turnansky)\
**Replies:** 2\
**Last updated:** [February 4, 2026, 5:31pm UTC](https://dev-discuss.pytorch.org/t/helion-inductor-integration/3302 "2026-02-04T17:31:19Z")

</div>

I am curious about the plan for Helion support in Inductor. Is there an RFC for this? For a given graph, is the intention to have the backend be purely triton vs purely Helion or is there a strategy to be able to mix an…

---

## [How to get graph weights in torch compile with custom backend?](https://dev-discuss.pytorch.org/t/how-to-get-graph-weights-in-torch-compile-with-custom-backend/3285)

<div class="topic-metadata">

**Author:** [@RunnerZhong](https://dev-discuss.pytorch.org/u/RunnerZhong)\
**Replies:** 1\
**Last updated:** [January 27, 2026, 10:53pm UTC](https://dev-discuss.pytorch.org/t/how-to-get-graph-weights-in-torch-compile-with-custom-backend/3285 "2026-01-27T22:53:57Z")

</div>

In custom backend fx graph, all weights or parameters are lifted to placeholder type. And how can I get the weights/parameters origin value in this FX graph ? I want to do something in custom backend during compile.

---

## [TorchDynamo: An Experiment in Dynamic Python Bytecode Transformation](https://dev-discuss.pytorch.org/t/torchdynamo-an-experiment-in-dynamic-python-bytecode-transformation/361)

<div class="topic-metadata">

**Author:** [@jansel](https://dev-discuss.pytorch.org/u/jansel)\
**Replies:** 10\
**Last updated:** [December 12, 2025, 4:54am UTC](https://dev-discuss.pytorch.org/t/torchdynamo-an-experiment-in-dynamic-python-bytecode-transformation/361 "2025-12-12T04:54:43Z")

</div>

In Next Steps for PyTorch Compilers, we laid out a vision of deploying eager mode PyTorch to more production settings and investing in using compilers to make eager mode faster and easier to maintain. This move away …

---

## [Question regarding horizontal fusion](https://dev-discuss.pytorch.org/t/question-regarding-horizontal-fusion/3275)

<div class="topic-metadata">

**Author:** [@Pawl](https://dev-discuss.pytorch.org/u/Pawl)\
**Replies:** 2\
**Last updated:** [December 10, 2025, 1:48pm UTC](https://dev-discuss.pytorch.org/t/question-regarding-horizontal-fusion/3275 "2025-12-10T13:48:17Z")

</div>

I am trying to get my head around Inductor. From inductor/fx\_passes/ and inductor/codegen it seems like there is support for both vertical and horizontal fusion. By vertical fusion I guess, what is meant is kernel fusion…

---

## [Torch.compile support for Python 3.14 completed](https://dev-discuss.pytorch.org/t/torch-compile-support-for-python-3-14-completed/3276)

<div class="topic-metadata">

**Author:** [@williamwen42](https://dev-discuss.pytorch.org/u/williamwen42)\
**Replies:** 0\
**Last updated:** [December 5, 2025, 10:28pm UTC](https://dev-discuss.pytorch.org/t/torch-compile-support-for-python-3-14-completed/3276 "2025-12-05T22:28:57Z")

</div>

Signalboosting that torch.compile is now compatible with Python 3.14! You can try it out today with the nightly PyTorch binaries. 3.14 support will be included in the next PyTorch release, 2.10. The no-GIL builds, 3.13t…

---

## [Min-cut optimal(\*) recomputation (i.e. activation checkpointing) with AOTAutograd](https://dev-discuss.pytorch.org/t/min-cut-optimal-recomputation-i-e-activation-checkpointing-with-aotautograd/467)

<div class="topic-metadata">

**Author:** [@Chillee](https://dev-discuss.pytorch.org/u/Chillee)\
**Replies:** 11\
**Last updated:** [December 3, 2025, 9:05pm UTC](https://dev-discuss.pytorch.org/t/min-cut-optimal-recomputation-i-e-activation-checkpointing-with-aotautograd/467 "2025-12-03T21:05:47Z")

</div>

TL;DR: We’ve implemented a min-cut based recomputation pass with AOTAutograd + NVFuser that consistently improves both memory and runtime across a wide range of models (including the TorchBench suite) for GPU training. I…

---

## [Could I static link a model using pytorch as an elf?](https://dev-discuss.pytorch.org/t/could-i-static-link-a-model-using-pytorch-as-an-elf/3250)

<div class="topic-metadata">

**Author:** [@ZenusZhang](https://dev-discuss.pytorch.org/u/ZenusZhang)\
**Replies:** 1\
**Last updated:** [October 11, 2025, 6:26am UTC](https://dev-discuss.pytorch.org/t/could-i-static-link-a-model-using-pytorch-as-an-elf/3250 "2025-10-11T06:26:16Z")

</div>

Hello everyone. I want to write aten kernel for a new riscv-vector cpu which haven’t taped out and only have scmodel. The cost of dynamic link is fairly high both on time and space, which is hard to be handled. I’ve kn…

---

## [Hy are AOTI packaged models faster than torch.compile models with the same inputs?](https://dev-discuss.pytorch.org/t/hy-are-aoti-packaged-models-faster-than-torch-compile-models-with-the-same-inputs/3240)

<div class="topic-metadata">

**Author:** [@skv66](https://dev-discuss.pytorch.org/u/skv66)\
**Replies:** 2\
**Last updated:** [October 9, 2025, 5:24am UTC](https://dev-discuss.pytorch.org/t/hy-are-aoti-packaged-models-faster-than-torch-compile-models-with-the-same-inputs/3240 "2025-10-09T05:24:30Z")

</div>

Hi all, I’ve been experimenting with both torch.compile and AOTI compiled packaged models. Here’s my setup: Same model architectures (e.g., ResNet18, etc.) Same input shapes (e.g., \[1, 3, 224, 224\]) Both run i…

---

## [A new strategy for automatic custom operators functionalization](https://dev-discuss.pytorch.org/t/a-new-strategy-for-automatic-custom-operators-functionalization/2733)

<div class="topic-metadata">

**Author:** [@laithsakka](https://dev-discuss.pytorch.org/u/laithsakka)\
**Replies:** 3\
**Last updated:** [September 28, 2025, 1:47am UTC](https://dev-discuss.pytorch.org/t/a-new-strategy-for-automatic-custom-operators-functionalization/2733 "2025-09-28T01:47:53Z")

</div>

A new strategy for automatic custom operators functionalization that enables re-inplacing on view tensors. with @zou3519 TLDR We shipped a new auto functionalization strategy for custom operators that automatically and …

---

## [Disabling Codegen-Specific Fusions in TorchInductor for Per-Op Kernel Generation](https://dev-discuss.pytorch.org/t/disabling-codegen-specific-fusions-in-torchinductor-for-per-op-kernel-generation/3226)

<div class="topic-metadata">

**Author:** [@skv66](https://dev-discuss.pytorch.org/u/skv66)\
**Replies:** 3\
**Last updated:** [September 4, 2025, 2:32am UTC](https://dev-discuss.pytorch.org/t/disabling-codegen-specific-fusions-in-torchinductor-for-per-op-kernel-generation/3226 "2025-09-04T02:32:36Z")

</div>

Hi all, I’m currently working on benchmarking TorchInductor’s optimization impact by selectively disabling various fusion and scheduling passes. So far, I have: Disabled inlining optimizations (Previous Post) Disabled…

---

## [How to turn off inlining / force materialization in TorchInductor during torch.compile?](https://dev-discuss.pytorch.org/t/how-to-turn-off-inlining-force-materialization-in-torchinductor-during-torch-compile/3198)

<div class="topic-metadata">

**Author:** [@skv66](https://dev-discuss.pytorch.org/u/skv66)\
**Replies:** 4\
**Last updated:** [August 29, 2025, 10:36pm UTC](https://dev-discuss.pytorch.org/t/how-to-turn-off-inlining-force-materialization-in-torchinductor-during-torch-compile/3198 "2025-08-29T22:36:16Z")

</div>

Hi all, I’m trying to reproduce the “Without inlining” ablation from the PyTorch 2.0 paper (the table that reports speedups for “All TorchInductor optimizations”, and then “Without inlining”, “Without fusion”, and “With…

---

## [📢 PyTorch Compiler H2, 25 Roadmap Review & Q&A — Join the Conversation!](https://dev-discuss.pytorch.org/t/pytorch-compiler-h2-25-roadmap-review-q-a-join-the-conversation/3222)

<div class="topic-metadata">

**Author:** [@miladm](https://dev-discuss.pytorch.org/u/miladm)\
**Replies:** 0\
**Last updated:** [August 27, 2025, 10:31pm UTC](https://dev-discuss.pytorch.org/t/pytorch-compiler-h2-25-roadmap-review-q-a-join-the-conversation/3222 "2025-08-27T22:31:57Z")

</div>

Hello PyTorch Community! :waving\_hand: We’re excited to announce an Open Source Roadmap OSS Q&A session for PyTorch Compile. This will be your opportunity to hear directly from the roadmap leads, ask questions, and shar…

---

## [Confusion About LibTorch Binary Naming and ABI on "Get Started" Page](https://dev-discuss.pytorch.org/t/confusion-about-libtorch-binary-naming-and-abi-on-get-started-page/3220)

<div class="topic-metadata">

**Author:** [@rastna12](https://dev-discuss.pytorch.org/u/rastna12)\
**Replies:** 0\
**Last updated:** [August 22, 2025, 6:54pm UTC](https://dev-discuss.pytorch.org/t/confusion-about-libtorch-binary-naming-and-abi-on-get-started-page/3220 "2025-08-22T18:54:46Z")

</div>

Hey all, I’m confused by the LibTorch binary naming and ABI labeling on the PyTorch “Get Started” page (https://pytorch.org/get-started/locally/). When selecting Linux and LibTorch with CUDA 12.6, the page provides the …

---

## [Torch.export functionalization changed?](https://dev-discuss.pytorch.org/t/torch-export-functionalization-changed/3142)

<div class="topic-metadata">

**Author:** [@ssilva-encharge](https://dev-discuss.pytorch.org/u/ssilva-encharge)\
**Replies:** 2\
**Last updated:** [July 24, 2025, 5:40pm UTC](https://dev-discuss.pytorch.org/t/torch-export-functionalization-changed/3142 "2025-07-24T17:40:24Z")

</div>

Hey folks, I’m struggling to produce an OutputKind.BUFFER\_MUTATION in the signature of a torch.export’ed program. I swear this used to work. I took the exact example from the docs: If I run this script: import torch …

---

## [Export\_memory\_timeline doesn't work with torch.profile](https://dev-discuss.pytorch.org/t/export-memory-timeline-doesnt-work-with-torch-profile/3143)

<div class="topic-metadata">

**Author:** [@nimisha](https://dev-discuss.pytorch.org/u/nimisha)\
**Replies:** 0\
**Last updated:** [July 23, 2025, 6:21am UTC](https://dev-discuss.pytorch.org/t/export-memory-timeline-doesnt-work-with-torch-profile/3143 "2025-07-23T06:21:40Z")

</div>

When using PyTorch Profiler’s export\_memory\_timeline on a standard (eager-mode) model, the resulting HTML visualization displays memory allocations with distinct colors and labels—parameters and activations are clearly i…

---

## [Error during build setup, has anyone experienced this?](https://dev-discuss.pytorch.org/t/error-during-build-setup-has-anyone-experienced-this/3120)

<div class="topic-metadata">

**Author:** [@SRIRAMDADI1](https://dev-discuss.pytorch.org/u/SRIRAMDADI1)\
**Replies:** 1\
**Last updated:** [July 20, 2025, 10:17pm UTC](https://dev-discuss.pytorch.org/t/error-during-build-setup-has-anyone-experienced-this/3120 "2025-07-20T22:17:30Z")

</div>

The error message is below was wondering if anyone knew the cause: bin\\torch\_cpu.dll : fatal error LNK1120: 1 unresolved externals ninja: build stopped: subcommand failed. – Building version 2.8.0a0+git31e1274 — Tryi…

---

## [Extern declaration of the entity XXX is treated as a static definition](https://dev-discuss.pytorch.org/t/extern-declaration-of-the-entity-xxx-is-treated-as-a-static-definition/3112)

<div class="topic-metadata">

**Author:** [@Ind1x1](https://dev-discuss.pytorch.org/u/Ind1x1)\
**Replies:** 0\
**Last updated:** [July 10, 2025, 6:25am UTC](https://dev-discuss.pytorch.org/t/extern-declaration-of-the-entity-xxx-is-treated-as-a-static-definition/3112 "2025-07-10T06:25:34Z")

</div>

I attempted to add two global device pointers in aten/native/cuda/CUDALoops.cuh to support a certain functionality. These pointers are defined in another file, myKernel.cu, and declared as extern in myKernel.cuh. Howeve…

---

## [Does dynamo trigger real kernel execution?](https://dev-discuss.pytorch.org/t/does-dynamo-trigger-real-kernel-execution/3081)

<div class="topic-metadata">

**Author:** [@guoyejun](https://dev-discuss.pytorch.org/u/guoyejun)\
**Replies:** 4\
**Last updated:** [July 1, 2025, 1:56am UTC](https://dev-discuss.pytorch.org/t/does-dynamo-trigger-real-kernel-execution/3081 "2025-07-01T01:56:39Z")

</div>

I think the answer is NO according to Dynamo Overview — PyTorch 2.7 documentation Dynamo hooks into the frame evaluation API in CPython (PEP 523) to dynamically modify Python bytecode right before it is executed. It re…

---

## [Torch.compile support for Python 3.13 completed](https://dev-discuss.pytorch.org/t/torch-compile-support-for-python-3-13-completed/2738)

<div class="topic-metadata">

**Author:** [@williamwen42](https://dev-discuss.pytorch.org/u/williamwen42)\
**Replies:** 2\
**Last updated:** [June 27, 2025, 9:46pm UTC](https://dev-discuss.pytorch.org/t/torch-compile-support-for-python-3-13-completed/2738 "2025-06-27T21:46:54Z")

</div>

Signalboosting that torch.compile is now compatible with Python 3.13! You can try it out today with the nightly PyTorch binaries. 3.13 support will be included in the next PyTorch release, 2.6. (Note that we currently st…

---

## [\[question\] how dynamo or aot-autograd optimize slice operation?](https://dev-discuss.pytorch.org/t/question-how-dynamo-or-aot-autograd-optimize-slice-operation/3089)

<div class="topic-metadata">

**Author:** [@Yepgang](https://dev-discuss.pytorch.org/u/Yepgang)\
**Replies:** 1\
**Last updated:** [June 27, 2025, 8:35am UTC](https://dev-discuss.pytorch.org/t/question-how-dynamo-or-aot-autograd-optimize-slice-operation/3089 "2025-06-27T08:35:37Z")

</div>

For the following case, the fx graph obtained by the inductor entry no longer contains the slice op. Does anyone know where this part is optimized? e.g. slice ut case import torch class NaiveModel(torch.nn.Module): …

[Next page](https://dev-discuss.pytorch.org/c/compiler/5.md?page=1)
