# New Contributor Looking for Mentorship!

**URL:** <https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637>\
**Category:** Uncategorized\
**Created:** [December 4, 2024, 6:50pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637 "2024-12-04T18:50:49Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![EmmettBicker](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/emmettbicker/32/2413_2.png) [@EmmettBicker](https://dev-discuss.pytorch.org/u/EmmettBicker)\
**Post date:** [December 4, 2024, 6:50pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/1 "2024-12-04T18:50:49Z")

</div>

Hi everyone,

I’m Emmett, an eighteen-year-old who’s passionate about contributing to PyTorch! I’ve been self-teaching myself machine learning for the past three years and I took a gap year before entering college to do independent research [(this is my most recent ongoing project)](https://github.com/EmmettBicker/IDEAL). I want to build my skills as a machine learning developer and I received advice that contributing to open source projects would be the best method to do so. I very frequently use PyTorch so it seemed like an obvious choice to contribute.

I would like to spend the upcoming months contributing to PyTorch full-time to build my skills and would really appreciate guidance/mentorship :).

So far I’ve made these two prs:

> <https://github.com/pytorch/pytorch/pull/141908>
>
> Fixes #127352 
> 
> Adds a linter to check for \`# flake8: noqa\` comments that appe…ar mid-file. It's implemented by using the built-in module tokenize to scan all comments and from those comments it identifies \`# flake8: noqa\` comments with a regex. It only flags errors after the first non-comment token (so the \`# flake8: noqa\` can only be preceded by other comments).
> 
> There are a considerable amount of files have mid-file \`# flake8: noqa\` comments which you can see by running
> \`\`\`bash
> lintrunner lint --take FLAKE8\_COMMENT --all-files
> \`\`\`
> 
> I wasn't sure how to deal with these preexisting mid-file comments, as adding in this linter would then make calls to \`lintrunner --all-files\` fail where it was previously successful, so I'd greatly appreciate any advice on that issue before the PR is considered for merging.

> <https://github.com/pytorch/pytorch/pull/142049>
>
> Works on #141828
> 
> This commit adds to \`tools/pyi/gen\_pyi.py\` to generate more …type hints in \`\_nn.pyi\`. I mainly created the hints by looking at the corresponding .h file in build/atn/src/Aten/ops, but there were often inconsistencies in the header files (such as excluding an out parameter or naming the first input "self" when in python it's "input") so I made tests to validate my hints that I didn't include in the commit. An example of one of these tests is attached at the bottom of the pr.
> 
> I made a significant amount of progress, but there's still more to be done. Adding more stubs would be great, and some of the existing type stubs need to be fixed as I tested a few and they were missing an out parameter in their stub or the first input parameter was named "self" in the stub but the actual function's was "input".
> 
> I added the hints for cross\_entropy\_loss, binary\_cross\_entropy, l1\_loss, smooth\_l1\_loss, max\_pool2d\_with\_indices, max\_pool3d\_with\_indices, huber\_loss, mse\_loss, multilabel\_margin\_loss, soft\_margin\_loss, multi\_margin\_loss, upsample\_nearest1d, upsample\_nearest2d, upsample\_nearest3d, \_upsample\_nearest\_exact1d, \_upsample\_nearest\_exact2d, and \_upsample\_nearest\_exact3d,
> 
> but 
> 
> max\_unpool3d, adaptive\_avg\_pool2d, adaptive\_avg\_pool3d, glu, relu6\_, relu6, elu, hardsigmoid\_, silu\_, silu, mish\_, mish, hardswish\_, hardswish, nll\_loss\_nd, upsample\_linear1d, \_upsample\_bilinear2d\_aa, upsample\_bilinear2d, upsample\_trilinear3d, \_upsample\_bicubic2d\_aa, upsample\_bicubic2d, im2col, col2im
> 
> all still need hints.
> 
> Here's an example of the tests I didn't include:
> \`\`\`python
> def test\_cross\_entropy\_loss():
> """
> cross\_entropy\_loss": \[
> "def cross\_entropy\_loss({}) -\> Tensor: ...".
> format(
> ", ".join(
> \[
> "input: Tensor",
> "target: Tensor",
> "weight: Optional\[Tensor\] = None",
> "reduction: int = 1",
> "ignore\_index: int = -100",
> "label\_smoothing: float = 0.0",
> \]
> )
> )
> """
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(InvalidType(), InvalidType())
> raise Exception # these make sure the above code raises an error
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'input' (position 1) must be Tensor, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'target' (position 2) must be Tensor, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'weight' (position 3) must be Tensor, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, t, InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'reduction' (position 4) must be int, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, t, i, InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'ignore\_index' (position 5) must be int, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, t, i, i, InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss(): argument 'label\_smoothing' (position 6) must be float, not InvalidType", e.\_\_str\_\_()
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, t, i, i, f)
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "ignore\_index is not supported for floating point target", e.\_\_str\_\_() # means it ran the code
> 
> try:
> torch.\_C.\_nn.cross\_entropy\_loss(t, t, t, i, i, f, out=InvalidType())
> raise Exception
> except Exception as e:
> assert e.\_\_str\_\_() == "cross\_entropy\_loss() got an unexpected keyword argument 'out'", e.\_\_str\_\_() 
> 
> \`\`\`

And I’m interested in pretty much all tasks, especially more involved ones that require me to understand PyTorch more in depth.

I’m especially interested in working on the following issues:

> <https://github.com/pytorch/pytorch/issues/118153>
>
> \### 🚀 The feature, motivation and pitch
> 
> Currently, \`aten::view\_as\_real\` is not …implemented for sparse tensors. A good reason to implement this is to enable us to support SparseAdam with complex dtypes.
> 
> \### Alternatives
> 
> \_No response\_
> 
> \### Additional context
> 
> \_No response\_
> 
> cc @alexsamardzic @nikitaved @pearu @cpuhrsch @amjames @bhosmer @jcaip @vincentqb @jbschlosser @albanD @crcrpar

> <https://github.com/pytorch/pytorch/issues/140881>
>
> \### 🚀 The feature, motivation and pitch
> 
> https://github.com/apple/ml-cross-entro…py Better version of cross entropy for large vocab sizes, this allows small models to have a dramatically higher batch size than otherwise. Implementing this conditionally would be trivial based on the input / output dimension and dtype of the input. There is a Pure triton implementation available that we can use in a compiler pass or add a dynamic dispatch for the faster version for A100s
> 
> \### Alternatives
> 
> \_No response\_
> 
> \### Additional context
> 
> \_No response\_
> 
> cc @albanD

---

<div class="post-metadata">

**Author:** ![janeyx99](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/janeyx99/32/352_2.png) [@janeyx99](https://dev-discuss.pytorch.org/u/janeyx99)\
**Post date:** [December 9, 2024, 9:12pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/2 "2024-12-09T21:12:30Z")

</div>

Hi! Nice to see this after engaging with you in recent discussions like [Optimizers' `differentiable` flag doesn't work · Issue #141832 · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/issues/141832)! Since you’re interested in contributing more than just piecewise, I’d be excited to work with you in tackling a more holistic medium-sized project around differentiable optimizers if you’re interested.

You’ve already gotten some context around action items, so this is to set a clearer goal to frame the bigger picture. Ultimately, we want to see differentiable optimizers have better support, test coverage, and documentation that they do today.

What would an ideal end state look like?  
(1) Better support: people can run differentiable optimizers with lr, betas, and weight\_decay as Tensors that require grad, meaning people can train their optimizer hyper params.

(2) Better documentation: we have a tutorial in [GitHub - pytorch/tutorials: PyTorch tutorials.](https://github.com/pytorch/tutorials) showing a real use case for differentiable optimizers, and our pytorch/pytorch documentation has a simpler code example. We also raise proper errors/warnings within the code linking to these resources.

(3) Fuller test coverage: our differentiable tests were excluded from our general test infrastructure migration to OptimizerInfo, but ideally, we’d use OptimizerInfos for these tests as well. An example test case that we should have our differentiable tests look like could be found in [pytorch/test/test\_optim.py at main · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/blob/main/test/test_optim.py#L913), see `test_foreach_large_tensor`. We’d want to use the OptimizerInfo infra to encompass all the new tests we want to add, like lr as a Tensor, etc.

Like most destinations, this end state can be achieve from several directions, and here’s a sample path taken from what we already delineated in the linked issue above:  
Step 0 (could be done in parallel or first, based on preference): Migrate the current differentiable tests [pytorch/test/optim/test\_optim.py at main · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/blob/main/test/optim/test_optim.py#L52) to use OptimizerInfos + expand test coverage.  
Step 1: support tensor LR when differentiable is True for SGD. Add a test case and docs in the code.  
Step 2: now what if the tensor LR requires grad? Make sure this works and add a test case and docs in the code.  
Step 3: Expand the above to different optimizers, Adam, AdamW, Adagrad, etc. Of course, add test cases and corresponding docs. This might be when it’d be good to consider using OptimizerInfos if you haven’t yet.  
Step 4: Add error messaging.  
Step 5: Add overarching docs on how to use differentiable optimizers and what’s supported. I could also see this being step 1, with gradual improvements as steps 1-3 are completed.

Let me know what you think!

---

<div class="post-metadata">

**Author:** ![EmmettBicker](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/emmettbicker/32/2413_2.png) [@EmmettBicker](https://dev-discuss.pytorch.org/u/EmmettBicker)\
**Post date:** [December 10, 2024, 3:01am UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/3 "2024-12-10T03:01:21Z")

</div>

This sounds wonderful and I would absolutely love to work with you on this project! Working on broader differentiable optimizer supports seems really meaningful and exciting. I would love to start working on step 0 before starting step 1, instead of working on both in parallel, because I haven’t used OptimizerInfo and I think I’ll have a bit to learn.

Thank you so much for this opportunity! I’ll start tomorrow after I tie up loose ends with other PyTorch PRs. Do you have a preferred method of communication for us to use?

---

<div class="post-metadata">

**Author:** ![janeyx99](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/janeyx99/32/352_2.png) [@janeyx99](https://dev-discuss.pytorch.org/u/janeyx99)\
**Post date:** [December 10, 2024, 7:00pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/4 "2024-12-10T19:00:10Z")

</div>

Cool! I’ve sent you a message (I think? I haven’t sent messages on dev discuss before lol) for further communication.

By the way, when I meant in parallel, I just meant it could be done any time, not necessarily at the same time. For example, feel free to do step 1 independently, and then look at step 0. I find that it is easiest to start with something you’re already a bit familiar with, and there’s quite a lot of flexibility here in where you want to start.

---

<div class="post-metadata">

**Author:** ![EmmettBicker](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/emmettbicker/32/2413_2.png) [@EmmettBicker](https://dev-discuss.pytorch.org/u/EmmettBicker)\
**Post date:** [December 10, 2024, 7:18pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/5 "2024-12-10T19:18:41Z")

</div>

Ok that makes sense! And I have received your message.

---

<div class="post-metadata">

**Author:** ![dscamiss](https://yyz2.discourse-cdn.com/flex036/user_avatar/dev-discuss.pytorch.org/dscamiss/32/2665_2.png) [@dscamiss](https://dev-discuss.pytorch.org/u/dscamiss)\
**Post date:** [April 17, 2025, 9:30pm UTC](https://dev-discuss.pytorch.org/t/new-contributor-looking-for-mentorship/2637/6 "2025-04-17T21:30:33Z")

</div>

Hi all, are you looking for another contributor on this project? I’m also trying to build up PyTorch experience and I think I can help out. What’s the current status? From what I can tell, it looks like “Step 1” above (tensor `lr` for SGD) is complete but might be missing test code.
