# Torch.ao.quantization Migration Plan

**URL:** <https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810>\
**Category:** Uncategorized\
**Created:** [February 25, 2025, 1:05am UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810 "2025-02-25T01:05:56Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jerryzh168](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/jerryzh168/32/1740_2.png) [@jerryzh168](https://dev-discuss.pytorch.org/u/jerryzh168)\
**Post date:** [February 25, 2025, 1:05am UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/1 "2025-02-25T01:05:56Z")

</div>

# Goal

The goal for the doc is to lay out the plan for deprecating and migrating quantization flows in torch.ao.quantization.

Note: This is a follow up to [Clarification of PyTorch Quantization Flow Support (in pytorch and torchao)](https://dev-discuss.pytorch.org/t/clarification-of-pytorch-quantization-flow-support-in-pytorch-and-torchao/2809) to clarify our migration plan for `torch.ao.quantization`.

# What is in `torch.ao.quantization`

| Flow | Release Status | Features | Backends | Note |
| --- | --- | --- | --- | --- |
| Eager Mode Quantization | beta | post training static, dynamic and weight only quantization, and quantization aware training (for static quantization), and numeric debugging tool. | x86 (fbgemm) and ARM CPU (qnnpack) | Quantized operators are using quantized Tensor in C++, that we plan to deprecate |
| TorchScript Graph Mode Quantization | prototype | post training static and dynamic quantization | x86 (fbgemm) and ARM CPU (qnnpack) | Quantized operators are using quantized Tensor in C++, that we plan to deprecate |
| FX Graph Mode Quantization | prototype | Post training static, dynamic, weight only, QAT, numeric suite | X86 (fbgemm/onednn) and ARM CPU (qnnpack/xnnpack) | Quantized operators are using quantized Tensor in C++, that we plan to deprecate |
| PT2E Quantization | prototype | Post Training static, dynamic, weight only, QAT, numeric debugger | X86 (onednn), ARM CPU (xnnpack), and many other mobile devices (boltnn, qualcomm, apple, turing, jarvis etc.) | Using pytorch native Tensors |
| | | | | |

# Flow Support in 2024

For some data points in terms of support, in 2024,

- We have fixed the following issues for QAT and PTQ, QAT is onboarding new customers in H1 2024, most of the fixes are for supporting new use cases, PTQ fixes are mostly bug fixes or making things more general.

- We did not receive or fix any issues related to fx, eager or torchscript quantization as far as I know

# Proposed Support Status

Overall I think we can have the following two statuses:

- Long Term Support
  - We commit to support the flow long term
  - We commit to fulfilling important feature request from other teams
  - We commit to bug fixes

- Phasing Out
  - We won’t add new features
  - We only commit to critical bug fixes

 ![Screenshot 2025-02-24 at 13.55.16](https://canada1.discourse-cdn.com/flex036/uploads/pytorch1/original/2X/6/6d6d200729cbf6e0b3d476c3fcaf032336b1ce00.png)

# Proposed Action Items

For PT2E Quantization, I think it would be better if we move the the implementation to torchao.

 ![Screenshot 2025-02-24 at 13.55.52](https://canada1.discourse-cdn.com/flex036/uploads/pytorch1/original/2X/a/aed2ac2e8e083e5ef6f53e21cc139ccb6b8cd678.png)

For other workflows, I think we can keep them in pytorch for now, we can revisit the plan for deleting code if the usage drops to a certain point.

In terms of how we do the migration, what we agreed on in torchao meeting is the following:

1. For code that is used by eager and fx mode quantization like observers and fake\_quant modules, we can keep these in pytorch/pytorch and import them from torchao

2. [1-2 weeks] For pt2e flow related code, we plan to duplicate them in torchao repository, new development with happen in torchao repository and older code in pytorch is kept for BC purposes

3. After we replicate pt2e flow code in torchao, we’ll also ask people to migrate to torchao APIs

- [2 weeks] Internally, torchao team will take care of changing API imports in fbcode
- [2 weeks] Externally, we can add a warning in torch.ao.quantization saying this will be deprecated soon, and potentially delete the pt2e code in 1-2 releases, we’ll add deprecation warning for all other workflows as well that have “Phase Out” support status

1. [not related to migration] We can also have new development such as adding groupwise observer in torchao after we have duplicated the pt2e flow code in torchao

We can target the above to be done by the end of H1 2025.

---

<div class="post-metadata">

**Author:** ![maajidkhann](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/maajidkhann/32/1772_2.png) [@maajidkhann](https://dev-discuss.pytorch.org/u/maajidkhann)\
**Post date:** [March 13, 2025, 6:07am UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/2 "2025-03-13T06:07:21Z")

</div>

@jerryzh168 As per [Torch.ao.quantization Migration Plan](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810),  
PT2E Quantization will be Long Term support moving further from PyTorch team. Right now, we have validated PT2E on ARM CPU’s and it only leverages FP32 kernels for compute. While this helps reducing memory footprint but the performance takes a hit as the compute happens still in FP32 and there are additional overheads.

[(prototype) PyTorch 2 Export Post Training Quantization — PyTorch Tutorials 2.6.0+cu124 documentation](https://pytorch.org/tutorials/prototype/pt2e_quant_ptq.html) also confirms in PT2E, the weights are still in fp32 right now and that you might do constant propagation for quantize op to get integer weights in the future.

Can we know the plans and details when Meta will introduce INT8 inference with PT2E. As this is the recommended quantization flow and the others are set for deprecation, it is very critical to leverage INT8 weights using PT2E.

---

<div class="post-metadata">

**Author:** ![jerryzh168](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/jerryzh168/32/1740_2.png) [@jerryzh168](https://dev-discuss.pytorch.org/u/jerryzh168)\
**Post date:** [March 14, 2025, 9:18pm UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/3 "2025-03-14T21:18:46Z")

</div>

int8 weights and int8 ops can be supported through torch.compile today, please take a look at [PyTorch 2 Export Quantization with X86 Backend through Inductor — PyTorch Tutorials 2.6.0+cu124 documentation](https://pytorch.org/tutorials/prototype/pt2e_quant_x86_inductor.html) as an example, relevant int8 ops for x86 backend can be found in [pytorch/aten/src/ATen/native/quantized/library.cpp at a0893475ba91f2b5c71b31af2be6b716b584ce48 · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/blob/a0893475ba91f2b5c71b31af2be6b716b584ce48/aten/src/ATen/native/quantized/library.cpp#L251), where we take int8/uint8 pytorch Tensors along with quantization parameters and output int8/uint8 Tensors.

convert\_pt2e folds the quantize op with fp32 weights by default, so should see ”int8 weights → dequant → linear → quant" pattern by default (and then you can fuse the pattern to int8 op) I think, let me know if it does not work.

---

<div class="post-metadata">

**Author:** ![anzr299](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/anzr299/32/3125_2.png) [@anzr299](https://dev-discuss.pytorch.org/u/anzr299)\
**Post date:** [January 27, 2026, 9:43am UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/4 "2026-01-27T09:43:11Z")

</div>

Hi @jerryzh168 will FakeQuantize ([https://github.com/pytorch/pytorch/blob/fbdc163dde6ead830d622bb7a13174091ebe4c53/torch/ao/quantizat…](https://github.com/pytorch/pytorch/blob/fbdc163dde6ead830d622bb7a13174091ebe4c53/torch/ao/quantization/fake_quantize.py#L70 "https://github.com/pytorch/pytorch/blob/fbdc163dde6ead830d622bb7a13174091ebe4c53/torch/ao/quantization/fake\_quantize.py#L70")) be moved to torchao as well? From my understanding of the migration plans, FakeQuantize, Observers etc. should be removed in the current or following release. Is that right?

---

<div class="post-metadata">

**Author:** ![jerryzh168](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/jerryzh168/32/1740_2.png) [@jerryzh168](https://dev-discuss.pytorch.org/u/jerryzh168)\
**Post date:** [January 27, 2026, 8:56pm UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/5 "2026-01-27T20:56:05Z")

</div>

FakeQuantize is actually duplicated in torchao right now

also it is used by eager / fx flow as well so we can’t really remove them from pytorch

---

<div class="post-metadata">

**Author:** ![anzr299](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/anzr299/32/3125_2.png) [@anzr299](https://dev-discuss.pytorch.org/u/anzr299)\
**Post date:** [January 28, 2026, 6:39am UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/6 "2026-01-28T06:39:25Z")

</div>

Thank you for the response!

I meant to ask if it will be removed from torch.ao.

what about completely removing it from torch.ao?

---

<div class="post-metadata">

**Author:** ![jerryzh168](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/jerryzh168/32/1740_2.png) [@jerryzh168](https://dev-discuss.pytorch.org/u/jerryzh168)\
**Post date:** [January 28, 2026, 7:08pm UTC](https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810/7 "2026-01-28T19:08:37Z")

</div>

hmmm, we can’t really remove right now because of Meta internal product using these APIs, we’re still thinking about the longer term deprecation / deletion plan. We may work together with TorchScript to deprecate these together in the future.
