UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible

Abduslam F A Nwesri; Nabila A S Shinbir; Hassan Ebrahem

doi:10.18653/v1/2023.arabicnlp-1.64

UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible

Abduslam F A Nwesri, Nabila A S Shinbir, Hassan Ebrahem

Correct Metadata for

Use this form to create a GitHub issue with structured data describing the correction. You will need a GitHub account. Once you create that issue, the correction will be reviewed by a staff member.

⚠️ Mobile Users: Submitting this form to create a new issue will only work with github.com, not the GitHub Mobile app.

Important: The Anthology treat PDFs as authoritative. Please use this form only to correct data that is out of line with the PDF. See our corrections guidelines if you need to change the PDF.

Title Adjust the title. Retain tags such as <fixed-case>.

Authors Adjust author names and order to match the PDF.

Abstract Correct abstract if needed. Retain XML formatting tags such as <tex-math>. You may use <b>...</b> for bold, <i>...</i> for italic, and <url>...</url> for URLs.

Verification against PDF Ensure that the new title/authors match the snapshot below. (If there is no snapshot or it is too small, consult the PDF.)

Authors concatenated from the text boxes above:

ALL author names match the snapshot above—including middle initials, hyphens, and accents.

Abstract

In this paper we present our approach towards Arabic Dialect identification which was part of the Fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). We tested several techniques to identify Arabic dialects. We obtained the best result by fine-tuning the pre-trained MARBERTv2 model with a modified training dataset. The training set was expanded by sorting tweets based on dialects, concatenating every two adjacent tweets, and adding them to the original dataset as new tweets. We achieved 82.87 on F1 score and we were at the seventh position among 16 participants.

Anthology ID:: 2023.arabicnlp-1.64
Volume:: Proceedings of ArabicNLP 2023
Month:: December
Year:: 2023
Address:: Singapore (Hybrid)
Editors:: Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, Rawan Almatham
Venues:: ArabicNLP | WS
SIG:: SIGARAB
Publisher:: Association for Computational Linguistics
Note:
Pages:: 620–624
Language:
URL:: https://aclanthology.org/2023.arabicnlp-1.64/
DOI:: 10.18653/v1/2023.arabicnlp-1.64
Bibkey:
Cite (ACL):: Abduslam F A Nwesri, Nabila A S Shinbir, and Hassan Ebrahem. 2023. UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible. In Proceedings of ArabicNLP 2023, pages 620–624, Singapore (Hybrid). Association for Computational Linguistics.
Cite (Informal):: UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible (Nwesri et al., ArabicNLP 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.arabicnlp-1.64.pdf
Video:: https://aclanthology.org/2023.arabicnlp-1.64.mp4

PDF Cite Search Video Fix data

Export citation

BibTeX
MODS XML
Endnote
Preformatted

@inproceedings{nwesri-etal-2023-uot,
    title = "{U}o{T} at {NADI} 2023 shared task: Automatic {A}rabic Dialect Identification is Made Possible",
    author = "Nwesri, Abduslam F A  and
      Shinbir, Nabila A S  and
      Ebrahem, Hassan",
    editor = "Sawaf, Hassan  and
      El-Beltagy, Samhaa  and
      Zaghouani, Wajdi  and
      Magdy, Walid  and
      Abdelali, Ahmed  and
      Tomeh, Nadi  and
      Abu Farha, Ibrahim  and
      Habash, Nizar  and
      Khalifa, Salam  and
      Keleg, Amr  and
      Haddad, Hatem  and
      Zitouni, Imed  and
      Mrini, Khalil  and
      Almatham, Rawan",
    booktitle = "Proceedings of ArabicNLP 2023",
    month = dec,
    year = "2023",
    address = "Singapore (Hybrid)",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.arabicnlp-1.64/",
    doi = "10.18653/v1/2023.arabicnlp-1.64",
    pages = "620--624",
    abstract = "In this paper we present our approach towards Arabic Dialect identification which was part of the Fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). We tested several techniques to identify Arabic dialects. We obtained the best result by fine-tuning the pre-trained MARBERTv2 model with a modified training dataset. The training set was expanded by sorting tweets based on dialects, concatenating every two adjacent tweets, and adding them to the original dataset as new tweets. We achieved 82.87 on F1 score and we were at the seventh position among 16 participants."
}

Download as File

<?xml version="1.0" encoding="UTF-8"?>
<modsCollection xmlns="http://www.loc.gov/mods/v3">
<mods ID="nwesri-etal-2023-uot">
    <titleInfo>
        <title>UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible</title>
    </titleInfo>
    <name type="personal">
        <namePart type="given">Abduslam</namePart>
        <namePart type="given">F</namePart>
        <namePart type="given">A</namePart>
        <namePart type="family">Nwesri</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <name type="personal">
        <namePart type="given">Nabila</namePart>
        <namePart type="given">A</namePart>
        <namePart type="given">S</namePart>
        <namePart type="family">Shinbir</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <name type="personal">
        <namePart type="given">Hassan</namePart>
        <namePart type="family">Ebrahem</namePart>
        <role>
            <roleTerm authority="marcrelator" type="text">author</roleTerm>
        </role>
    </name>
    <originInfo>
        <dateIssued>2023-12</dateIssued>
    </originInfo>
    <typeOfResource>text</typeOfResource>
    <relatedItem type="host">
        <titleInfo>
            <title>Proceedings of ArabicNLP 2023</title>
        </titleInfo>
        <name type="personal">
            <namePart type="given">Hassan</namePart>
            <namePart type="family">Sawaf</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Samhaa</namePart>
            <namePart type="family">El-Beltagy</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Wajdi</namePart>
            <namePart type="family">Zaghouani</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Walid</namePart>
            <namePart type="family">Magdy</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Ahmed</namePart>
            <namePart type="family">Abdelali</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Nadi</namePart>
            <namePart type="family">Tomeh</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Ibrahim</namePart>
            <namePart type="family">Abu Farha</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Nizar</namePart>
            <namePart type="family">Habash</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Salam</namePart>
            <namePart type="family">Khalifa</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Amr</namePart>
            <namePart type="family">Keleg</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Hatem</namePart>
            <namePart type="family">Haddad</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Imed</namePart>
            <namePart type="family">Zitouni</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Khalil</namePart>
            <namePart type="family">Mrini</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <name type="personal">
            <namePart type="given">Rawan</namePart>
            <namePart type="family">Almatham</namePart>
            <role>
                <roleTerm authority="marcrelator" type="text">editor</roleTerm>
            </role>
        </name>
        <originInfo>
            <publisher>Association for Computational Linguistics</publisher>
            <place>
                <placeTerm type="text">Singapore (Hybrid)</placeTerm>
            </place>
        </originInfo>
        <genre authority="marcgt">conference publication</genre>
    </relatedItem>
    <abstract>In this paper we present our approach towards Arabic Dialect identification which was part of the Fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). We tested several techniques to identify Arabic dialects. We obtained the best result by fine-tuning the pre-trained MARBERTv2 model with a modified training dataset. The training set was expanded by sorting tweets based on dialects, concatenating every two adjacent tweets, and adding them to the original dataset as new tweets. We achieved 82.87 on F1 score and we were at the seventh position among 16 participants.</abstract>
    <identifier type="citekey">nwesri-etal-2023-uot</identifier>
    <identifier type="doi">10.18653/v1/2023.arabicnlp-1.64</identifier>
    <location>
        <url>https://aclanthology.org/2023.arabicnlp-1.64/</url>
    </location>
    <part>
        <date>2023-12</date>
        <extent unit="page">
            <start>620</start>
            <end>624</end>
        </extent>
    </part>
</mods>
</modsCollection>

Download as File

%0 Conference Proceedings
%T UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible
%A Nwesri, Abduslam F. A.
%A Shinbir, Nabila A. S.
%A Ebrahem, Hassan
%Y Sawaf, Hassan
%Y El-Beltagy, Samhaa
%Y Zaghouani, Wajdi
%Y Magdy, Walid
%Y Abdelali, Ahmed
%Y Tomeh, Nadi
%Y Abu Farha, Ibrahim
%Y Habash, Nizar
%Y Khalifa, Salam
%Y Keleg, Amr
%Y Haddad, Hatem
%Y Zitouni, Imed
%Y Mrini, Khalil
%Y Almatham, Rawan
%S Proceedings of ArabicNLP 2023
%D 2023
%8 December
%I Association for Computational Linguistics
%C Singapore (Hybrid)
%F nwesri-etal-2023-uot
%X In this paper we present our approach towards Arabic Dialect identification which was part of the Fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). We tested several techniques to identify Arabic dialects. We obtained the best result by fine-tuning the pre-trained MARBERTv2 model with a modified training dataset. The training set was expanded by sorting tweets based on dialects, concatenating every two adjacent tweets, and adding them to the original dataset as new tweets. We achieved 82.87 on F1 score and we were at the seventh position among 16 participants.
%R 10.18653/v1/2023.arabicnlp-1.64
%U https://aclanthology.org/2023.arabicnlp-1.64/
%U https://doi.org/10.18653/v1/2023.arabicnlp-1.64
%P 620-624

Download as File

Markdown (Informal)

[UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible](https://aclanthology.org/2023.arabicnlp-1.64/) (Nwesri et al., ArabicNLP 2023)

UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible (Nwesri et al., ArabicNLP 2023)

ACL

Abduslam F A Nwesri, Nabila A S Shinbir, and Hassan Ebrahem. 2023. UoT at NADI 2023 shared task: Automatic Arabic Dialect Identification is Made Possible. In Proceedings of ArabicNLP 2023, pages 620–624, Singapore (Hybrid). Association for Computational Linguistics.