Keyphrase Generation for Scientific Document Retrieval

Florian Boudin; Ygor Gallina; Akiko Aizawa

doi:10.18653/v1/2020.acl-main.105

Keyphrase Generation for Scientific Document Retrieval

Florian Boudin, Ygor Gallina, Akiko Aizawa

Correct Metadata for

Use this form to create a GitHub issue with structured data describing the correction. You will need a GitHub account. Once you create that issue, the correction will be reviewed by a staff member.

⚠️ Mobile Users: Submitting this form to create a new issue will only work with github.com, not the GitHub Mobile app.

Important: The Anthology treat PDFs as authoritative. Please use this form only to correct data that is out of line with the PDF. See our corrections guidelines if you need to change the PDF.

Title Adjust the title. Retain tags such as <fixed-case>.

Authors Adjust author names and order to match the PDF.

Abstract Correct abstract if needed. Retain XML formatting tags such as <tex-math>. You may use <b>...</b> for bold, <i>...</i> for italic, and <url>...</url> for URLs.

Verification against PDF Ensure that the new title/authors match the snapshot below. (If there is no snapshot or it is too small, consult the PDF.)

Authors concatenated from the text boxes above:

ALL author names match the snapshot above—including middle initials, hyphens, and accents.

Abstract

Sequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval. This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models. Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in text- keyphrases, and generalizing models across domains. Our code is available at https://github.com/boudinfl/ir-using-kg

Anthology ID:: 2020.acl-main.105
Volume:: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
Month:: July
Year:: 2020
Address:: Online
Editors:: Dan Jurafsky, Joyce Chai, Natalie Schluter, Joel Tetreault
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1118–1126
Language:
URL:: https://aclanthology.org/2020.acl-main.105/
DOI:: 10.18653/v1/2020.acl-main.105
Bibkey:
Cite (ACL):: Florian Boudin, Ygor Gallina, and Akiko Aizawa. 2020. Keyphrase Generation for Scientific Document Retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1118–1126, Online. Association for Computational Linguistics.
Cite (Informal):: Keyphrase Generation for Scientific Document Retrieval (Boudin et al., ACL 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.acl-main.105.pdf
Video:: http://slideslive.com/38928736

PDF Cite Search Video Fix data

Export citation

BibTeX
MODS XML
Endnote
Preformatted

@inproceedings{boudin-etal-2020-keyphrase,
    title = "Keyphrase Generation for Scientific Document Retrieval",
    author = "Boudin, Florian  and
      Gallina, Ygor  and
      Aizawa, Akiko",
    editor = "Jurafsky, Dan  and
      Chai, Joyce  and
      Schluter, Natalie  and
      Tetreault, Joel",
    booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
    month = jul,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2020.acl-main.105/",
    doi = "10.18653/v1/2020.acl-main.105",
    pages = "1118--1126",
    abstract = "Sequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval. This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models. Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in text- keyphrases, and generalizing models across domains. Our code is available at \url{https://github.com/boudinfl/ir-using-kg}"
}

Download as File

<?xml version="1.0" encoding="UTF-8"?>
<modsCollection xmlns="http://www.loc.gov/mods/v3">
<mods ID="boudin-etal-2020-keyphrase">
    <titleInfo>
        <title>Keyphrase Generation for Scientific Document Retrieval</title>
    </titleInfo>
    <name type="personal">
        <namePart type="given">Florian</namePart>
        <namePart type="family">Boudin</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <name type="personal">
        <namePart type="given">Ygor</namePart>
        <namePart type="family">Gallina</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <name type="personal">
        <namePart type="given">Akiko</namePart>
        <namePart type="family">Aizawa</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <originInfo>
        <dateIssued>2020-07</dateIssued>
    </originInfo>
    <typeOfResource>text</typeOfResource>
    <relatedItem type="host">
        <titleInfo>
            <title>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</title>
        </titleInfo>
        <name type="personal">
            <namePart type="given">Dan</namePart>
            <namePart type="family">Jurafsky</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Joyce</namePart>
            <namePart type="family">Chai</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Natalie</namePart>
            <namePart type="family">Schluter</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Joel</namePart>
            <namePart type="family">Tetreault</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <originInfo>
            <publisher>Association for Computational Linguistics</publisher>
            <place>
                <placeTerm type="text">Online</placeTerm>
            </place>
        </originInfo>
        <genre authority="marcgt">conference publication</genre>
    </relatedItem>
    <abstract>Sequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval. This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models. Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in text- keyphrases, and generalizing models across domains. Our code is available at https://github.com/boudinfl/ir-using-kg</abstract>
    <identifier type="citekey">boudin-etal-2020-keyphrase</identifier>
    <identifier type="doi">10.18653/v1/2020.acl-main.105</identifier>
    <location>
        <url>https://aclanthology.org/2020.acl-main.105/</url>
    </location>
    <part>
        <date>2020-07</date>
        <extent unit="page">
            <start>1118</start>
            <end>1126</end>
        </extent>
    </part>
</mods>
</modsCollection>

Download as File

%0 Conference Proceedings
%T Keyphrase Generation for Scientific Document Retrieval
%A Boudin, Florian
%A Gallina, Ygor
%A Aizawa, Akiko
%Y Jurafsky, Dan
%Y Chai, Joyce
%Y Schluter, Natalie
%Y Tetreault, Joel
%S Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
%D 2020
%8 July
%I Association for Computational Linguistics
%C Online
%F boudin-etal-2020-keyphrase
%X Sequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval. This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models. Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in text- keyphrases, and generalizing models across domains. Our code is available at https://github.com/boudinfl/ir-using-kg
%R 10.18653/v1/2020.acl-main.105
%U https://aclanthology.org/2020.acl-main.105/
%U https://doi.org/10.18653/v1/2020.acl-main.105
%P 1118-1126

Download as File

Markdown (Informal)

[Keyphrase Generation for Scientific Document Retrieval](https://aclanthology.org/2020.acl-main.105/) (Boudin et al., ACL 2020)

Keyphrase Generation for Scientific Document Retrieval (Boudin et al., ACL 2020)

ACL

Florian Boudin, Ygor Gallina, and Akiko Aizawa. 2020. Keyphrase Generation for Scientific Document Retrieval. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1118–1126, Online. Association for Computational Linguistics.