TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing

Pooja Guhan, Uttaran Bhattacharya, Somdeb Sarkhel, Vahid Azizi, Xiang Chen, Saayan Mitra, Aniket Bera, Dinesh Manocha


Abstract
Given a source and its edited version performed based on human instructions in natural language, how do we extract the underlying edit operations, to automatically replicate similar edits on other images? This is the problem of reverse designing, and we present TAME-RD, a model to solve this problem. TAME-RD automatically learns from the complex interplay of image editing operations and the natural language instructions to learn fully specified edit operations. It predicts both the underlying image edit operations as discrete categories and their corresponding parameter values in the continuous space.We accomplish this by mapping together the contextual information from the natural language text and the structural differences between the corresponding source and edited images using the concept of pre-post effect. We demonstrate the efficiency of our network through quantitative evaluations on multiple datasets. We observe improvements of 6-10% on various accuracy metrics and 1.01X-4X on the RMSE score and the concordance correlation coefficient for the corresponding parameter values on the benchmark GIER dataset. We also introduce I-MAD, a new two-part dataset: I-MAD-Dense, a collection of approximately 100K source and edited images, together with automatically generated text instructions and annotated edit operations, and I-MAD-Pro, consisting of about 1.6K source and edited images, together with text instructions and annotated edit operations provided by professional editors. On our dataset, we observe absolute improvements of 1-10% on the accuracy metrics and 1.14X–5X on the RMSE score.
Anthology ID:
2024.findings-acl.637
Volume:
Findings of the Association for Computational Linguistics ACL 2024
Month:
August
Year:
2024
Address:
Bangkok, Thailand and virtual meeting
Editors:
Lun-Wei Ku, Andre Martins, Vivek Srikumar
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
10710–10727
Language:
URL:
https://aclanthology.org/2024.findings-acl.637
DOI:
Bibkey:
Cite (ACL):
Pooja Guhan, Uttaran Bhattacharya, Somdeb Sarkhel, Vahid Azizi, Xiang Chen, Saayan Mitra, Aniket Bera, and Dinesh Manocha. 2024. TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing. In Findings of the Association for Computational Linguistics ACL 2024, pages 10710–10727, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
Cite (Informal):
TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing (Guhan et al., Findings 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.findings-acl.637.pdf