# Wav2vec2.0 model support in torchaudio

**URL:** <https://dev-discuss.pytorch.org/t/wav2vec2-0-model-support-in-torchaudio/220>\
**Category:** Uncategorized\
**Created:** [May 13, 2021, 8:47pm UTC](https://dev-discuss.pytorch.org/t/wav2vec2-0-model-support-in-torchaudio/220 "2021-05-13T20:47:48Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthrok](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/mthrok/32/122_2.png) [@mthrok](https://dev-discuss.pytorch.org/u/mthrok)\
**Post date:** [May 13, 2021, 8:47pm UTC](https://dev-discuss.pytorch.org/t/wav2vec2-0-model-support-in-torchaudio/220/1 "2021-05-13T20:47:48Z")

</div>

Hi PyTorch dev community

I posted the project about supporting wav2vec2.0 in torchaudio.

> <https://github.com/pytorch/audio/issues/1506>
>
> \# wav2vec2.0 model with TorchScript support
> 
> Following the popularity of wav2v…ec2.0 models, we plan to add wav2vec2.0 model in torchaudio.
> 
> \## What are we going to add?
> 
> 1. Model definitions in \`torchaudio\`, that supports TorchScript, specifically \`torch.jit.script\`.
> 2. Conversion of pretrained models from Hugging Face's \`transformers\` (and \`fairseq\`).\*
> 
> \\\* For importing models of these libraries, you will need to install them, but they will be optional dependencies of \`torchaudio\`.
> 
> \## What is the value of this addition?
> 
> With the support for TorchScript, one can deploy wav2vec2.0 models into non-Python environments, such as C++ applications and mobiles.
> 
> There is \[an example of running wav2vec2.0 in mobile\](https://github.com/pytorch/android-demo-app/tree/master/SpeechRecognition), but due to the limitation on TorchScript support, it is constrained on the fixed number of time frames. We want to improve this by adding a model definition that supports scripting method.
> 
> \## So is this inference only?
> 
> At the beginning, our focus and the intended use is inference. However, since the model is authored with \`torch.nn.Module\`, you can still plug-in it to your training as well. Note that the initial version will behave differently in training mode. (Things like random layer drop in Transformer or weight normalization will not be included.)
> 
> For training, we recommend using \`fairseq\`, and Hugging Face \`transformers\`, which are confirmed to work.
> 
> \## Is there a prototype?
> 
> You can see the example of running wav2vec2.0 pre-trained models in \[the prototype branch\](https://github.com/pytorch/audio/tree/wav2vec2-prototype/examples/libtorchaudio/speech\_recognition).
> 
> \## What are the accuracies?
> 
> With the prototype above, we ran models in C++ and evaluated them on Librispeech.
> 
> For \[models from \`fairseq\` repository\](https://github.com/pytorch/fairseq/tree/master/examples/wav2vec), we also ran them in Python, and got the exact same results.
> 
> | Architecture | Fine Tune | test-clean | test-other |
> | ----------------------------------------- | ----------:|-----------:| ----------:|
> | Base \<br/\>\<code\>wav2vec\_small\_960h\</code\> | 960h | 3.1 | 7.7 |
> | Large \<br/\>\<code\>wav2vec\_big\_960h\</code\> | 960h | 2.6 | 5.9 |
> | Large (LV-60) \<br/\>\<code\>wav2vec2\_vox\_960h\_new\</code\> | 960h | 2.9 | 6.2 |
> | Large (LV-60) + Self Training \<br/\>\<code\>(wav2vec\_vox\_960h\_pl\</code\> | 960h | 1.9 | 4.5 |
> 
> For models from Hugging Face \`transformers\`, we got a similar result.
> 
> | Architecture | Fine Tune | test-clean | test-other |
> | ----------------------------------------- | ----------:|-----------:| ----------:|
> | Base \<br/\>\<code\>facebook/wav2vec2-base-960h\</code\> | 960h | 3.1 | 7.8 |
> | Large \<br/\>\<code\>facebook/wav2vec2-large-960h\</code\> | 960h | 2.5 | 5.8 |
> | Large (LV-60) \<br/\>\<code\>facebook/wav2vec2-large-960h-lv60\</code\> | 960h | 3.0 | 6.3 |
> | Large (LV-60) + Self Training \<br/\>\<code\>facebook/wav2vec2-large-960h-lv60-self\</code\> | 960h | 1.9 | 4.5 |
> 
> We also verified that models fine-tuned using Hugging Face \`transformers\` work. We ran \`facebook/wav2vec2-large-xlsr-53-german\` model on VoxForge Germany dataset and got WER of around 10.8.
> 
> \## What are the performances?
> 
> We compared the run time performaces of the original \`fairseq\` model and our C++ version. C++ version is slightly faster but, not significantly so. Quantization of the model might help this.
> 
> !\[Encoding RTF\](https://user-images.githubusercontent.com/855818/118184401-4cbb9800-b409-11eb-8180-bf56cfa15211.png)
> 
> \## What are the future works?
> 
> There are couple of things we can think of for the future direction. (But please note that we do not promise any future outcome.)
> 
> 1. Addition of TorchScript-able CTC decoder in \`torchaudio\`.
> 2. Quantization support and mobile optimization support.
> 3. Training support.

If this interests you, please give us a feedback.

Thanks!
