PyMT5: multi-mode translation of natural language and Python code with transformers

Colin Clement; Dawn Drain; Jonathan Timcheck; Alexey Svyatkovskiy; Neel Sundaresan

doi:10.18653/v1/2020.emnlp-main.728

PyMT5: multi-mode translation of natural language and Python code with transformers

Colin Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, Neel Sundaresan

Abstract

Simultaneously modeling source code and natural language has many exciting applications in automated software development and understanding. Pursuant to achieving such technology, we introduce PyMT5, the Python method text-to-text transfer transformer, which is trained to translate between all pairs of Python method feature combinations: a single model that can both predict whole methods from natural language documentation strings (docstrings) and summarize code into docstrings of any common style. We present an analysis and modeling effort of a large-scale parallel corpus of 26 million Python methods and 7.7 million method-docstring pairs, demonstrating that for docstring and method generation, PyMT5 outperforms similarly-sized auto-regressive language models (GPT2) which were English pre-trained or randomly initialized. On the CodeSearchNet test set, our best model predicts 92.1% syntactically correct method bodies, achieved a BLEU score of 8.59 for method generation and 16.3 for docstring generation (summarization), and achieved a ROUGE-L F-score of 24.8 for method generation and 36.7 for docstring generation.

Anthology ID:: 2020.emnlp-main.728
Volume:: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Month:: November
Year:: 2020
Address:: Online
Editors:: Bonnie Webber, Trevor Cohn, Yulan He, Yang Liu
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 9052–9065
Language:
URL:: https://aclanthology.org/2020.emnlp-main.728/
DOI:: 10.18653/v1/2020.emnlp-main.728
Bibkey:
Cite (ACL):: Colin Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020. PyMT5: multi-mode translation of natural language and Python code with transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9052–9065, Online. Association for Computational Linguistics.
Cite (Informal):: PyMT5: multi-mode translation of natural language and Python code with transformers (Clement et al., EMNLP 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.emnlp-main.728.pdf
Video:: https://slideslive.com/38939242

PDF Cite Search Video Fix data