# Summary of Issues Found in Adapting Models to use TorchScript

**URL:** <https://dev-discuss.pytorch.org/t/summary-of-issues-found-in-adapting-models-to-use-torchscript/154>\
**Category:** compiler\
**Created:** [February 26, 2021, 7:29pm UTC](https://dev-discuss.pytorch.org/t/summary-of-issues-found-in-adapting-models-to-use-torchscript/154 "2021-02-26T19:29:55Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kevin.stephano](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/kevin.stephano/32/105_2.png) [@kevin.stephano](https://dev-discuss.pytorch.org/u/kevin.stephano)\
**Post date:** [February 26, 2021, 7:29pm UTC](https://dev-discuss.pytorch.org/t/summary-of-issues-found-in-adapting-models-to-use-torchscript/154/1 "2021-02-26T19:29:55Z")

</div>

This is a [Google Doc](https://docs.google.com/document/d/1rJqfUKbuOzlZ8RXWRUBOKPHhKaf1f28LV1piTREL1u4/edit?usp=sharing) that summarizes and links the discussions found in the [Adapting Models to use TorchScript and Getting them to Produce Fusions](https://dev-discuss.pytorch.org/t/adapting-models-to-use-torchscript-and-getting-them-to-produce-fusions/129) post.

I am linking a doc because the top post appears not to be edit-able after a certain amount of time.

---

<div class="post-metadata">

**Author:** ![kevin.stephano](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/kevin.stephano/32/105_2.png) [@kevin.stephano](https://dev-discuss.pytorch.org/u/kevin.stephano)\
**Post date:** [March 16, 2021, 5:18pm UTC](https://dev-discuss.pytorch.org/t/summary-of-issues-found-in-adapting-models-to-use-torchscript/154/2 "2021-03-16T17:18:48Z")

</div>

We recently found a new issue that is also filed as [Pytorch Issue 54040](https://github.com/pytorch/pytorch/issues/54040) where TorchScript’s Autodiff is not respecting the `requires_grad` option of a tensor when calculating gradients such that unused gradients are unnecessarily calculated. The summary was updated. This was seen, in particular, on the mask applied to multihead attention in NLP networks.
