Natural language processing

Technology

For more than 40 years, NTT has been at the forefront of research and development of technologies that enable computers to understand and generate natural language used by humans. The applications of these technologies are diverse, including automatic dialogue, question answering, document summarization, translation, and information retrieval. In recent years, NTT has focused on methods that use large-scale language models, which are neural networks learned from large text corpora. We are also challenging the technology of "Vision-and-Language," a fusion of vision and language understanding. Our research and development are about understanding human communication in a multimodal way and generating it, with language at the core.

NTT Human Informatics Laboratories aim to realize a society that uses AI to understand the world and grow with others.

Research

  1. Understanding and generating natural language (under construction)
  2. Understanding and generating human communication in a multimodal (under construction)

Publications

2026

Conference Papers

  1. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi and Ryo Masumura, "Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2026), January 2026.
  2. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Takuhiro Kaneko, Shota Orihashi and Ryo Masumura, "Distribution Highlighted Reference-based Label Distribution Learning for Facial Age Estimation", in Proceedings of Winter Conference on Applications of Computer Vision (WACV 2026), March 2026.
  3. Yui Oka, Kentaro Hanafusa, Taku Hasegawa, Kyosuke Nishida and Kuniko Saito, "Probing Rotary Position Embeddings through Frequency Entropy", in Proceedings of The 14th International Conference on Learning Representations, April 2026.
  4. Yui Oka, Itsumi Saito, Kyosuke Nishida and Kuniko Saito, "Frequency Bands in RoPE: Base Frequency and Context Length Shape the Interpolation–Extrapolation Trade-off", in Proceedings of the 14th International Conference on Learning Representations (ICLR 2026), April 2026.
  5. Tomohiro Tanaka, Ryo Masumura, Naoki Makishima, Mana Ihori, Naotaka Kawata, Shota Orihashi, Satoshi Suzuki and Taiga Yamane, "Joint Autoregressive Modeling of Multi-talker Overlapped Speech Recognition and Translation", in Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026), May 2026.
  6. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Suzuka Yamada, Taiga Yamane, Naotaka Kawata and Satoshi Suzuki, "Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview", in Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026), May 2026.
  7. Kazutoshi Shinoda, Kosuke Nishida and Kyosuke Nishida, "Debiasing Reward Models via Causally Motivated Inference-Time Intervention", in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), July 2026.
  8. Shota Orihashi, Taiga Yamane, Naoki Makishima, Mana Ihori, Satoshi Suzuki, Tomohiro Tanaka and Ryo Masumura, "Social Group Activity Recognition from Still Images Using Conditional Token Sequence Generation", in Proceedings of International Conference on Image Processing (ICIP 2026), September 2026.
  9. Haruka Kawasaki, Ryota Tanaka and Kyosuke Nishida, "Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding", in Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR workshop), June 2026.
  10. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Yoshihiro Yamazaki, Keita Suzuki, Hiroaki Sugiyama and Kuniko Saito, "Let’s Put Ourselves in Sally’s Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models", in Proceedings of Findings of the Association for Computational Linguistics (EACL 2026), March 2026.

Papers

  1. Daiki Shiono, Shumpei Miyawaki, Ryota Tanaka and Jun Suzuki, "Instruction-Following Evaluation of Large Vision-Language Models", in New Generation Computing, January 2026.
  2. Ryunosuke Baba, Junya Morita, Ryuichiro Higashinaka and Yugo Takeuchi, "Modeling the Depth of Grounding through a Generative Cognitive Model", in Referential Communication Computers in Human Behavior: Artificial Humans, March 2026

2025

Conference Papers

  1. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Satoshi Suzuki, Saki Mizuno and Nobukatsu Hojo, "Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2025), February 2025.
  2. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama and Kuniko Saito, "ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2025), February 2025.
  3. Nobukatsu Hojo, Kazutoshi Shinoda, Yoshihiro Yamazaki, Keita Suzuki, Hiroaki Sugiyama, Kyosuke Nishida and Kuniko Saito, "GenerativeGUI: Dynamic GUI Generation leveraging LLMs for Enhanced User Interaction on Chat Interfaces", in Proceedings of Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI 2025), April 2025.
  4. Yui Oka, Taku Hasegawa, Kyosuke Nishida and Kuniko Saito, "Wavelet-based Positional Representation for Long Context", in Proceedings of the 13th International Conference on Learning Representations (ICLR 2025), April 2025.
  5. Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "Retrieval-Augmented Generation over Visually-Rich Documents", in Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025), June 2025.
  6. Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki,Shota Orihashi and Ryo Masumura, "SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-talker Overlapped ASR and Speaker Diarization", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), August 2025.
  7. Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi and Ryo Masumura, "Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-talker Overlapped Speech and Video", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), August 2025.
  8. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Taiga Yamane, Naotaka Kawata, Satoshi Suzuki and Taichi Katayama, "Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition", in Proceedings of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025), October 2025.
  9. Taiga Yamane, Ryo Masumura, Satoshi Suzuki and Shota Orihashi, "MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost", in Proceedings of International Conference on Computer Vision (ICCV 2025), October 2025.
  10. Tomohiro Tanaka, Ryo Masumura, Naoki Makishima, Mana Ihori, Shota Orihashi, Satoshi Suzuki and Taiga Yamane, "Semi-Supervised End-to-End Speech-to-Text Translation with Joint Text-to-Text and Speech-to-Text Decoding", in Proceedings of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025), October 2025.
  11. Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima and Naotaka Kawata, "MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection", in Proceedings of British Machine Vision Conference (BMVC 2025), November 2025.
  12. Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi and Ryo Masumura, "Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model", in Proceedings of Automatic Speech Recognition and Understanding Workshop (ASRU 2025), December 2025.
  13. Sora Kadotani, Kosuke Nishida and Kyosuke Nishida, "Learning from Hallucinations: Mitigating Hallucinations in LLMs via Internal Representation Intervention", in Proceedings of Findings of the International Joint Conference on Natural Language Processing & Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL Findings), December 2025.
  14. Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Satoshi Suzuki and Takafumi Moriya, "Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASR", in Proceedings of Automatic Speech Recognition and Understanding Workshop (ASRU 2025), December 2025.
  15. Ryunosuke Baba, Junya Morita, Takeru Amaya, Ryuichiro Higashinaka and Yugo Takeuchi, "Common Ground Building through Generative Cognitive Modules: Examining the Roles of Initial Perception, Imaging and Captionin", in Proceedings of Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2025), August 2025.
  16. Yosuke Ujigwa, Asuka Shiotani, Masato Takizawa, Eisuke Midorikawa, Ryuichiro Higashinaka and Kazunori Takashio, "Modality Analysis in Building Common Ground: Insights from a Collaborative Scene Reordering Task", in Proceedings of International Workshop on Spoken Dialogue Systems Technology (IWSDS 2025), May 2025.

Papers

  1. Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno and Ryo Masumura, "Data Stream-pairwise Bottleneck Transformer for Engagement Estimation from Video Conversation", Frontiers in Artificial Intelligence, vol.8, June 2025.

Tutorials

  1. Ryota Tanaka, Kosuke Nishida and Kyosuke Nishida, "Recent Advances in Large Language Models and Vision-and-Language Models Tutorials", in the 21st International Conference on Advanced Data Mining and Applications (ADMA 2025), October 2025.

2024

Conference Papers

  1. Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki, Jun Suzuki and Keisuke Sakaguchi, "Empirical Analysis of Large Vision Language Model against Goal Hijack via Visual Prompt Injection", in Proceedings of ACL Student Research Workshop (ASC SRW 2024), 2024.
  2. Haruto Yoshida, Ryota Tanaka, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "Does Vision Models Encode the Properties of Diagrams?", in Proceedings of ACL Student Research Workshop (ASC SRW 2024), 2024.
  3. Junya Morita, Tatsuya Yui, Takeru Amaya, Ryuichiro Higashinaka and Yugo Takeuchi, "Cognitive Architecture Toward Common Ground Sharing Among Humans and Generative AIs: Trial Modeling on Model-Model Interaction in Tangram Naming Task", in Proceedings of Proceedings of the AAAI Symposium Series (AAAI Symposium 2024), 2024.
  4. Ryo Masumura, Akihiko Takashima, Satoshi Suzuki and Shota Orihashi, "Born-Again Multi-Task Self-Training for Multi-Task Facial Emotion Recognition", in Proceedings of International Conference on Pattern Recognition (ICPR 2024), December 2024.
  5. Naotaka Kawata, Shota Orihashi; Satoshi Suzuki, Tomohiro Tanaka, Mana Ihori and Naoki Maikishima, "Block Refinement Learning for Improving Early Exit in Autoregressive ASR", in Proceedings of Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2024), December 2024.
  6. Kosuke Nishida, Kyosuke Nishida and Kuniko Saito, "Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes", in Proceedings of Conference on Empirical Methods in Natural Language aProcessing (EMNLP 2024), November 2024.
  7. Satoshi Suzuki, Shotaro Tora and Ryo Masumura, "Scene Generalized Multi-View Pedestrian Detection with Rotation-Based Augmentation and Regularization", in Proceedings of International Conference on Image Processing (ICIP 2024), October 2024.
  8. Taiga Yamane, Satoshi Suzuki, Ryo Masumura and Shotaro Tora, "MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with Transformer", in Proceedings of International Conference on Image Processing (ICIP 2024), October 2024.
  9. Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Satoshi Suzuki, Nobukatsu Hojo, Takafumi Moriya and Atsushi Ando, "Unified Multi-Talker ASR Modeling of with and without Target-speaker Enrollment", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  10. Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando and Ryo Masumura, "SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  11. Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno and Ryo Masumura, "Participant-Pair-Wise Bottleneck Transformer for Engagement Estimation from Video Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  12. Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa and Ryo Masumura, "Learning from Multiple Annotator Biased Labels in Multimodal Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (ICASSP 2024), September 2024.
  13. Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa and Ryo Masumura, "Learning from Multiple Annotator Biased Labels in Multimodal Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  14. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa and Ryo Masumura, "TALKING FACE GENERATION FOR IMPRESSION CONVERSION CONSIDERING SPEECH SEMANTICS", in Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024), April 2024.
  15. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa and Ryo Masumura, "Talking Face Generation for Impression Conversion Considering Speech Semantic", in Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024), April 2024.
  16. Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "InstructDoc: Zero-shot Generalization of Visual Document Understanding with Instructions", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2024), February 2024.

Papers

  1. Kazuki Sakai, Koh Mitsuda, Yuichiro Yoshikawa, Ryuichiro Higashinaka, Takashi Minato and Hiroshi Ishiguro, "Effects of Demonstrating Consensus Between Robots to Change User's Opinion", International Journal of Social Robotics 16(7), 2024.

Awards

  1. IEEE International Conference on Image Processing 2024 (ICIP2024) Best Industry Paper Award, Taiga Yamane, Satoshi Suzuki, Ryo Masumura and Shotaro Tora, "MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with Transformer"

2023

Conference Papers

  1. Keita Suzuki, Satoshi Suzuki, Ryo Masumura, Atsushi Ando and Naoki Makishima, "Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper Body", in Proceedings of the 5th ACM International Conference on Multimedia in Asia (MMAsia 2023), December 2023.
  2. Taku Hasegawa, Kyosuke Nishida, Koki Maeda and Kuniko Saito, "DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning", in Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), December 2023.
  3. Mihiro Uchida, Shota Orihashi, Akihiko Takashima, Yoshihiro Yamazaki and Ryo Masumura, "Open-Set Recognition for Facial-Expression Recognition", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  4. Satoshi Suzuki, Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando and Ryo Masumura, "ONDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training Framework", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  5. Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima and Ryo Masumura, "Distilling Knowledge of Bidirectional Language Model for Scene Text Recognition", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  6. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando and Ryo Masumura, "Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuray-Robustness Tradeoff", in Proceedings of International Conference on Computer Vision (ICCV 2023), October 2023.
  7. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka and Shota Orihashi, "Text-to-Text Pre-Training with Paraphrasing for Improving Transformer-based Image Captioning", in Proceedings of the 31st European Signal Processing Conference (EUSIPCO 2023), September 2023.
  8. Mana Ihori, Hiroshi Sato, Tomohiro Tanaka and Ryo Masumura, "Retrieval, Masking, and Generation: Feedback Comment Generation using Masked Comment Examples", in Proceedings of the 16th International Conference on Natural Language Generation (INLG 2023): Generation Challenge, September 2023.
  9. Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura, Saki Mizuno and Nobukatsu Hojo, "Transcribing Speech as Spoken and Written Dual Text Using an Autoregressive Model", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH), August 2023.
  10. Naoki Makishima, Keita Suzuki, Satoshi Suzuki, Atsushi Ando and Ryo Masumura, "Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  11. Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo and Atsushi Ando, "End-to-End Joint Target and Non-Target Speakers ASR", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  12. Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Hiroshi Sato and Tomohiro Tanaka, "Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  13. Saki Mizuno, Nobukatsu Hojo, Satoshi Kobashikawa and Ryo Masumura, "Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation", in Proceedings of the 48th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023), June 2023.
  14. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato, Taiga Yamane, Takanori Ashihara, Kohei Matsuura and Takafumi Moriya, "Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning", in Proceedings of the 48th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023), June 2023.
  15. Kosuke Nishida, Naoki Yoshinaga (Tokyo Univ.) and Kyosuke Nishida, "Self-Adaptive Named Entity Recognition by Retrieving Unstructured Knowledge", in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2023), May 2023.
  16. Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, Kuniko Saito, "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images", in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI 2023), February 2023.

2022

Conference Papers

  1. Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida and Sen Yoshida, "Japanese ASR-Robust Pre-trained Language Model with Pseudo-Error Sentences Generated by Grapheme-Phoneme Conversion", in Proceedings of the 2022 Conference of the International Speech Communication Association (INTERSPEECH 2022), pp. 2688-2692, September 2022.
  2. Fumio Nihei, Ryo Ishii, Yukiko Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama and Takao Nakamura, "Dialogue Acts Aided Important Utterance Detection Based on multiparty and multimodal information", in Proceedings of the 2022 Conference of the International Speech Communication Association (INTERSPEECH 2022), pp. 1086-1090, September 2022.
  3. Kosuke Nishida, Kyosuke Nishida, Shuichi Nishioka, "Improving Few-Shot Image Classification Using Machine- and User-Generated Natural Language Descriptions", in Proceedings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2022) (findings), pp. 1421-1430, July 2022. [arxiv]
  4. Shumpei Miyawaki (Tohoku Univ.), Taku Hasegawa, Kyosuke Nishida, Takuma Kato (Tohoku Univ.), Jun Suzuki (Tohoku Univ.), "Scene-Text Aware Image and Text Retrieval with Dual-Encoder", in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (ACL SRW 2022), pp. 422-433, May 2022.

2021

Conference Papers

  1. Kosuke Nishida, Kyosuke Nishida, Sen Yoshida, "Task-adaptive Pre-training of Language Models with Word Embedding Regularization", in Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021) (findings), pp. 4546-4553, August 2021. [arxiv]
  2. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Sen Yoshida: "Towards Interpretable and Reliable Reading Comprehension: A Pipeline Model with Unanswerability Prediction." in Proceddings of the 2021 International Joint Conference on Neural Networks ( IJCNN 2021), pp. 1-8, July 2021. [arxiv]
  3. Ryota Tanaka(*), Kyosuke Nishida(*), Sen Yoshida, "VisualMRC: Machine Reading Comprehension on Document Images", in Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI 2021), pp. 13878-13888, a Virtual Conference, February 2021. (*: equal contribution) (full paper, 1696/7911=21.4%) [arxiv] [project]

2020

Conference Papers

  1. Diana Galvan-Sosa (Tohoku Univ.), Jun Suzuki (Tohoku Univ.), Kyosuke Nishida, Koji Matsuda (Tohoku Univ.) and Kentaro Inui (Tohoku Univ.): Seeing the world through text: Evaluating image descriptions for commonsense reasoning in machine reading comprehension, in Proceedings of the Second Workshop on Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN 2020; in conjunction with COLING 2020), December 2020.
  2. Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda and Shoichiro Saito, "A Transformer-based Audio Captioning Model with Keyword Estimation", in Proceedings of the 21st Annual Conference of the International Speech Communication Association (INTERSPEECH 2020), October 2020. [arxiv]
  3. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Hisako Asano and Junji Tomita, "Unsupervised Domain Adaptation of Language Models for Reading Comprehension", in Proceedings of the 12th International Conference on Language Resources and Evaluation (LREC 2020), pp. 5392-5399, May 2020. [arxiv]

2019

Conference Papers

  1. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, and Takanobu Oba, "Generalized Large-Context Language Models based on Forward-Backward Hierarchical Encoder-Decoder Models", in Proceedings of the 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2019), pp.554-561, December 2019.
  2. Kyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda (Tokyo Univ.), Atsushi Otsuka, Hisako Asano and Junji Tomita, "Multi-style Generative Reading Comprehension", in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pp.2273-2284, July 2019. [arXiv] (long paper, 660/2905=22.7%)
  3. Kosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Itsumi Saito, Atushi Otuka, Hisako Asano and Junji Tomita, "Answering while Summarizing: Multi-task Learning for Multi-hop QA with Evidence Extraction", in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pp.2273-2284, 2335-2345, July 2019. [arXiv] (long paper, 660/2905=22.7%)
  4. Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida, Hisako Asano, and Junji Tomita, "A Simple but Effective Method to Incorporate Multi-turn Context with BERT for Conversational Machine Comprehension", in Proceedings of 1st Workshop on NLP for Conversational AI (NLP4ConvAI 2019; in conjunction with ACL 2019), Florence, Italy, July 2019. [arXiv]
  5. Atsushi Otsuka, Kyosuke Nishida, Itsumi Saito, Hisako Asano and Junji Tomita, "Specific Question Generation for Reading Comprehension", in Proceedings of the AAAI 2019 Reasoning for Complex QA (RCQA) Workshop, Honolulu, Hawaii, USA, January 2019.