# Intel GPU & CPU Enabling Status and Feature Plan – 2025 H1 Update

**URL:** <https://dev-discuss.pytorch.org/t/intel-gpu-cpu-enabling-status-and-feature-plan-2025-h1-update/2913>\
**Category:** hardware-backends\
**Created:** [April 15, 2025, 12:49pm UTC](https://dev-discuss.pytorch.org/t/intel-gpu-cpu-enabling-status-and-feature-plan-2025-h1-update/2913 "2025-04-15T12:49:53Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![EikanWang](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/eikanwang/32/992_2.png) [@EikanWang](https://dev-discuss.pytorch.org/u/EikanWang)\
**Post date:** [April 15, 2025, 12:49pm UTC](https://dev-discuss.pytorch.org/t/intel-gpu-cpu-enabling-status-and-feature-plan-2025-h1-update/2913/1 "2025-04-15T12:49:53Z")

</div>

## Intel GPU

### Context

We previously published the [Intel GPU Enabling Status and Feature Plan](https://dev-discuss.pytorch.org/t/intel-gpu-enabling-status-and-feature-plan/2389) to introduce Intel GPU support in PyTorch. Following the recent release of the [Meta PyTorch Team 2025 H1 Roadmaps](https://dev-discuss.pytorch.org/t/meta-pytorch-team-2025-h1-roadmaps/2794), the Intel PyTorch team has refreshed the status and feature plan to stay aligned with PyTorch community strategies, such as open-source adoption in head repositories (organic growth) through UX enhancements and foundational performance optimizations to drive sustained performance improvements.

**Recap**

Intel GPU in PyTorch is designed to provide a seamless GPU programming experience, covering both front-end and back-end integration. By leveraging Intel’s advancements in GPU technology, we enhance PyTorch’s performance and versatility, enabling significant workload acceleration and improved processing efficiency. This ensures an out-of-the-box experience for users on the Intel GPU platform while benefiting the broader PyTorch community.

**Areas Where We Excelled in Recent PyTorch Releases:**

- Over time expanded Intel GPU support to a wide range across Linux and Windows, including:

- Simple installation of torch-xpu PIP wheels and an effortless setup experience

- High ATen operation coverage with SYCL and oneDNN for smooth eager mode support with functionality and performance

- Notable speedups with `torch.compile` through default `TorchInductor` and `Triton` backend, proved by measurable performance gains with Hugging Face\*, TIMM, and TorchBench benchmarks

**Continued Efforts for Future PyTorch Releases:**

- Achieve PyTorch-native SOTA performance across important model categories and benchmarks on Intel GPUs

- Deliver a more streamlined and user-friendly experience on Intel GPUs

- Improve generalization of the PyTorch runtime and backend infrastructure to support a wide range of hardware backends

### Focus Areas for Continued Efforts

#### PyTorch Compiler Core

- PyTorch-native SOTA performance across key model categories and benchmarks, including generation models

- Exportability on Intel GPUs for deployment

- Coherent programming model across the board for the PyTorch Compiler (including both compile and export path) on Intel GPUs for better development UX

#### PyTorch Core Libraries

- Improve inference and training on Intel GPUs

- Reduce overall library entropy while maturing infrastructure to be more secure, performant, and stable on Intel GPUs

- Evolve Intel GPU support in PyTorch to embrace the emergent trend of test time computation.

- Architect Intel GPU support in `TorchCodec`

- Collaborate with the community to explore establishing a baseline OSS LLM for Intel GPU kernel generation

#### PyTorch Core Performance

- Attention Module on Intel GPUs

- INT8 Quantization - PT2E

- Support various performance and accuracy techniques in `torchao`, specifically

- Ensure `torchao` performs well on a diverse range of Intel GPU hardware for selected models such as Llama 3.1 8B and 3.2 11B:

- Contribute to `torchao` to establish a robust benchmarking infrastructure to track performance for inference on a set of key models on Intel GPUs

- Integrate `torchao` + XPU with SGLang, vLLM, HF Transformer, HF Diffuser

- Align fine-tuning UX with torchao and enable `torchao` on Intel GPUs in `torchtune`:

- Support MX dtypes in both PyTorch and `torchao` for Intel GPUs

#### PyTorch Distributed

- Native support for Intel® oneAPI Collective Communications Library (oneCCL) as `xccl` backend in `torch.distributed`

- Mature profiling and debugging tools for large scale-up and scale-out configurations

- Torch Titan: Enable XCCL support in torchtitan and showcase pretraining for Llama 3.1 LLM and reference MoE model

- Torch Tune: Enable XCCL distributed fine-tuning for standard Llama 3.x model, enable end-to-end distributed recipe(s) such as [torchtune/recipes/full\_finetune\_distributed.py at main · meta-pytorch/torchtune · GitHub](https://github.com/pytorch/torchtune/blob/main/recipes/full_finetune_distributed.py)

## Intel CPU

### Context

The Intel PyTorch team continues enhancing the performance and feature sets of PyTorch on Intel CPU platforms. In the past PyTorch releases, the Intel PyTorch team has:

- Optimized FlexAttention on Intel CPU platforms, which rapidly improves the performance of PyTorch on LLM for both first token latency and next token latency;
- Optimized BF16/FP16 to be Beta level for both eager mode and Inductor mode on Intel CPUs;
- Optimized TorchInductor CPP backend to be Beta level on Intel CPUs;

More can be found in past PyTorch release blogs.

### Focus Areas for Continued Efforts

- PyTorch-native SOTA performance of key LLM/Image-Generation models

- Competitive GEMM performance on Intel CPUs

- Mature FlexAttention to ensure reliability and high performance for different scenarios on Intel CPUs

- Reduce overall PyTorch Core/library entropy to be more stable and performant on Intel CPUs

## Summary

In this update, we summarized the progress we’ve made in recent PyTorch releases and outlined our continued efforts as a North Star to improve Intel GPU and Intel CPU support in future releases. Our achievements would not have been possible without the invaluable support from [Alban](https://github.com/albanD), [Andrey](https://github.com/atalman), [Bin](https://github.com/desertfire), [Jason](https://github.com/jansel), [Jerry](https://github.com/jerryzh168), and [Nikita](https://github.com/malfet) — we sincerely thank them for their contributions.

The Intel PyTorch team will continue to work closely with the PyTorch community to enhance Intel GPU and Intel GPU functionality and performance, covering eager mode, `torch.compile`, Distributed, and core libraries. At the same time, we are committed to contributing to the broader PyTorch ecosystem, helping make PyTorch increasingly friendly and extensible for integration with diverse hardware backends.

---

<div class="post-metadata">

**Author:** ![gottbrath](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/gottbrath/32/1043_2.png) [@gottbrath](https://dev-discuss.pytorch.org/u/gottbrath)\
**Post date:** [April 16, 2025, 8:05pm UTC](https://dev-discuss.pytorch.org/t/intel-gpu-cpu-enabling-status-and-feature-plan-2025-h1-update/2913/2 "2025-04-16T20:05:53Z")

</div>

Thanks a lot for sharing these updates!
