Does the agent say what it sees: an analysis in 3D question answering

Zhang, Haomeng

Does the agent say what it sees: an analysis in 3D question answering

Zhang, Haomeng

Permalink

https://hdl.handle.net/2142/120420

Description

Title

Does the agent say what it sees: an analysis in 3D question answering

Author(s)

Zhang, Haomeng

Issue Date

2023-04-27

Director of Research (if dissertation) or Advisor (if thesis)

Gui, Liangyan

Department of Study

Computer Science

Discipline

Computer Science

Degree Granting Institution

University of Illinois at Urbana-Champaign

Degree Name

M.S.

Degree Level

Thesis

Keyword(s)

Multi-modal Learning
3d Question Answering

Language

eng

Abstract

The emerging research topic of 3D Question and Answering involves complex challenges such as semantic language understanding, 3D scene comprehension, and establishing correspondence between language and target objects. In this thesis, I conduct a comprehensive analysis on the state-of-the-art 3D Question Answering benchmark model, ScanQA, by comparing the effectiveness of various auxiliary tasks, identifying limitations in current evaluation methods, and conducting additional ablation studies on the visual component. The findings and potential future directions for the 3D Question Answering task are discussed to assist researchers in their ongoing investigations.

Graduation Semester

2023-05

Type of Resource

Thesis

Handle URL

https://hdl.handle.net/2142/120420

Copyright and License Information

Owning Collections

Graduate Dissertations and Theses at Illinois PRIMARY

Graduate Theses and Dissertations at Illinois

Dissertations and Theses - Computer Science

Dissertations and Theses from the Siebel School of Computer Science

Does the agent say what it sees: an analysis in 3D question answering

Zhang, Haomeng

Permalink

Description

Owning Collections

Graduate Dissertations and Theses at Illinois PRIMARY

Dissertations and Theses - Computer Science

Log In