自然言語処理

技術カテゴリ紹介

NTTでは40年以上に渡って、対話・質問応答・文書要約・翻訳・情報検索など、人間が使う自然言語をコンピュータに理解・生成させるための技術の最前線で研究開発に取り組んで来ています。近年では、大量のテキストコーパスから学習されたニューラルネットワークである大規模言語モデルを用いた技術の発展に注力しています。さらには、「Vision-and-Language」と呼ばれる視覚と言語の融合理解をはじめ、言語を軸としたマルチモーダル理解・生成についても積極的に取り組んでいます。人間研ではこれらの取組を通じて、AIが人をとりまく世界を理解し人と相互に成長し合うような共生社会の実現をめざしています。

研究紹介

  1. 自然言語理解・生成技術
  2. マルチモーダル理解・生成技術

文献リスト

2026

国際会議

  1. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi and Ryo Masumura, "Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2026), January 2026.
  2. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Takuhiro Kaneko, Shota Orihashi and Ryo Masumura, "Distribution Highlighted Reference-based Label Distribution Learning for Facial Age Estimation", in Proceedings of Winter Conference on Applications of Computer Vision (WACV 2026), March 2026.
  3. Yui Oka, Kentaro Hanafusa, Taku Hasegawa, Kyosuke Nishida and Kuniko Saito, "Probing Rotary Position Embeddings through Frequency Entropy", in Proceedings of The 14th International Conference on Learning Representations, April 2026.
  4. Yui Oka, Itsumi Saito, Kyosuke Nishida and Kuniko Saito, "Frequency Bands in RoPE: Base Frequency and Context Length Shape the Interpolation–Extrapolation Trade-off", in Proceedings of the 14th International Conference on Learning Representations (ICLR 2026), April 2026.
  5. Tomohiro Tanaka, Ryo Masumura, Naoki Makishima, Mana Ihori, Naotaka Kawata, Shota Orihashi, Satoshi Suzuki and Taiga Yamane, "Joint Autoregressive Modeling of Multi-talker Overlapped Speech Recognition and Translation", in Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026), May 2026.
  6. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Suzuka Yamada, Taiga Yamane, Naotaka Kawata and Satoshi Suzuki, "Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview", in Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026), May 2026.
  7. Kazutoshi Shinoda, Kosuke Nishida and Kyosuke Nishida, "Debiasing Reward Models via Causally Motivated Inference-Time Intervention", in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), July 2026.
  8. Shota Orihashi, Taiga Yamane, Naoki Makishima, Mana Ihori, Satoshi Suzuki, Tomohiro Tanaka and Ryo Masumura, "Social Group Activity Recognition from Still Images Using Conditional Token Sequence Generation", in Proceedings of International Conference on Image Processing (ICIP 2026), September 2026.
  9. Haruka Kawasaki, Ryota Tanaka and Kyosuke Nishida, "Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding", in Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR workshop), June 2026.
  10. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Yoshihiro Yamazaki, Keita Suzuki, Hiroaki Sugiyama and Kuniko Saito, "Let’s Put Ourselves in Sally’s Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models", in Proceedings of Findings of the Association for Computational Linguistics (EACL 2026), March 2026.

論文

  1. Daiki Shiono, Shumpei Miyawaki, Ryota Tanaka and Jun Suzuki, "Instruction-Following Evaluation of Large Vision-Language Models", in New Generation Computing, January 2026.
  2. Ryunosuke Baba, Junya Morita, Ryuichiro Higashinaka and Yugo Takeuchi, "Modeling the Depth of Grounding through a Generative Cognitive Model", in Referential Communication Computers in Human Behavior: Artificial Humans, March 2026

招待・チュートリアル講演

  1. 折橋 翔太,"次世代メディア処理AI「MediaGnosis」における画像処理技術の研究と産業応用",画像電子学会セミナー Advanced Image Seminar 2026,2026年6月.
  2.  
  3. 風戸広史, “ソフトウェア開発に向けた⼤規模⾔語モデルの応⽤と発展", Developers Summit 2026, 2026年2月.

表彰

  1. 電子情報通信学会 IE賞:鈴木 聡志,山口 真弥,武田 翔一郎,山根 大河,牧島 直輝,庵 愛,田中 智大,折橋 翔太,増村 亮,"Vision-LanguageモデルのRobust Fine-tuningのための差分ベクトル等化",画像工学研究会,March 2026.
  2. 言語処理学会第32回年次大会(NLP2026) 優秀賞, 田中 涼太, 長谷川 拓, 西田 京介, "CMDR: 文脈を考慮したマルチモーダル文書検索"
  3. 言語処理学会第32回年次大会(NLP2026) 優秀賞, 吉田 遥音, 工藤 慧音, 青木 洋一, 田中 涼太, 斉藤 いつみ, 坂口 慶祐, 乾 健太郎, "大規模視覚言語モデル内部におけるダイアグラムの表現形成過程"
  4. 言語処理学会第32回年次大会(NLP2026) 若手奨励賞, 門谷 宙, "ハルシネーションから学ぶ:内部表現への介入によるハルシネーション抑制"
  5. 言語処理学会第32回年次大会(NLP2026) 優秀賞, 岡 佑依, 花房 健太郎, 長谷川 拓, 西田 京介, "周波数エントロピーによる位置埋込みの解明"
  6. 言語処理学会第32回年次大会(NLP2026) 優秀賞, 岡 佑依, 斉藤 いつみ, 西田 京介, "位置符号化の基底拡大戦略は外挿性能を制限する"

2025

国際会議

  1. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Satoshi Suzuki, Saki Mizuno and Nobukatsu Hojo, "Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2025), February 2025.
  2. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama and Kuniko Saito, "ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2025), February 2025.
  3. Nobukatsu Hojo, Kazutoshi Shinoda, Yoshihiro Yamazaki, Keita Suzuki, Hiroaki Sugiyama, Kyosuke Nishida and Kuniko Saito, "GenerativeGUI: Dynamic GUI Generation leveraging LLMs for Enhanced User Interaction on Chat Interfaces", in Proceedings of Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI 2025), April 2025.
  4. Yui Oka, Taku Hasegawa, Kyosuke Nishida and Kuniko Saito, "Wavelet-based Positional Representation for Long Context", in Proceedings of the 13th International Conference on Learning Representations (ICLR 2025), April 2025.
  5. Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "Retrieval-Augmented Generation over Visually-Rich Documents", in Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025), June 2025.
  6. Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki,Shota Orihashi and Ryo Masumura, "SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-talker Overlapped ASR and Speaker Diarization", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), August 2025.
  7. Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi and Ryo Masumura, "Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-talker Overlapped Speech and Video", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2025), August 2025.
  8. Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Taiga Yamane, Naotaka Kawata, Satoshi Suzuki and Taichi Katayama, "Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition", in Proceedings of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025), October 2025.
  9. Taiga Yamane, Ryo Masumura, Satoshi Suzuki and Shota Orihashi, "MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost", in Proceedings of International Conference on Computer Vision (ICCV 2025), October 2025.
  10. Tomohiro Tanaka, Ryo Masumura, Naoki Makishima, Mana Ihori, Shota Orihashi, Satoshi Suzuki and Taiga Yamane, "Semi-Supervised End-to-End Speech-to-Text Translation with Joint Text-to-Text and Speech-to-Text Decoding", in Proceedings of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025), October 2025.
  11. Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima and Naotaka Kawata, "MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection", in Proceedings of British Machine Vision Conference (BMVC 2025), November 2025.
  12. Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi and Ryo Masumura, "Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model", in Proceedings of Automatic Speech Recognition and Understanding Workshop (ASRU 2025), December 2025.
  13. Sora Kadotani, Kosuke Nishida and Kyosuke Nishida, "Learning from Hallucinations: Mitigating Hallucinations in LLMs via Internal Representation Intervention", in Proceedings of Findings of the International Joint Conference on Natural Language Processing & Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL Findings), December 2025.
  14. Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Satoshi Suzuki and Takafumi Moriya, "Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASR", in Proceedings of Automatic Speech Recognition and Understanding Workshop (ASRU 2025), December 2025.
  15. Ryunosuke Baba, Junya Morita, Takeru Amaya, Ryuichiro Higashinaka and Yugo Takeuchi, "Common Ground Building through Generative Cognitive Modules: Examining the Roles of Initial Perception, Imaging and Captionin", in Proceedings of Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2025), August 2025.
  16. Yosuke Ujigwa, Asuka Shiotani, Masato Takizawa, Eisuke Midorikawa, Ryuichiro Higashinaka and Kazunori Takashio, "Modality Analysis in Building Common Ground: Insights from a Collaborative Scene Reordering Task", in Proceedings of International Workshop on Spoken Dialogue Systems Technology (IWSDS 2025), May 2025.

論文

  1. Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno and Ryo Masumura, "Data Stream-pairwise Bottleneck Transformer for Engagement Estimation from Video Conversation", Frontiers in Artificial Intelligence, vol.8, June 2025.
  2. 岡 佑依, 柳本 大輝, 平尾 努, 西田 京介, "談話関係ラベル付き接続語認識", 自然言語処理, 2025年32巻2号, June 2025.

招待・チュートリアル講演

  1. 篠田一聡, “LLMは心の理論を持っているか?", 機械学習と数理モデルの融合と理論の深化 Ⅲ, 2025年10月.
  2. 西田京介, “NTT版大規模言語モデルtsuzumi 2について", LLM-jp勉強会, 2025年10月.
  3. 西田京介, “tsuzumi2: 進化したNTT版大規模言語モデル", weights and biases, Fully Connected Tokyo, 2025年10月.
  4. Ryota Tanaka, Kosuke Nishida and Kyosuke Nishida, "Recent Advances in Large Language Models and Vision-and-Language Models Tutorials", in the 21st International Conference on Advanced Data Mining and Applications (ADMA 2025), October 2025.
  5. 田中涼太, "VDocRAG: 視覚的文書に対する検索拡張生成 LLM-jp", マルチモーダルWG, 2025年3月.

表彰

  1. 言語処理学会第31回年次大会(NLP2025) 委員特別賞, 田中涼太, 壹岐太一, 長谷川拓, 西田京介, 齋藤邦子, 鈴木潤, "VDocRAG: 視覚的文書に対する検索拡張生成"
  2. 言語処理学会第31回年次大会(NLP2025) 優秀賞, 岡佑依, 長谷川拓, 西田京介, 齋藤邦子, "ウェーブレット位置符号化"
  3. 言語処理学会第31回年次大会(NLP2025) 優秀賞, 西田光甫, 西田京介, 齋藤邦子, "大規模言語モデルの再パラメタ化に基づく初期化による 損失スパイクの抑制"

2024

国際会議

  1. Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki, Jun Suzuki and Keisuke Sakaguchi, "Empirical Analysis of Large Vision Language Model against Goal Hijack via Visual Prompt Injection", in Proceedings of ACL Student Research Workshop (ASC SRW 2024), 2024.
  2. Haruto Yoshida, Ryota Tanaka, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "Does Vision Models Encode the Properties of Diagrams?", in Proceedings of ACL Student Research Workshop (ASC SRW 2024), 2024.
  3. Junya Morita, Tatsuya Yui, Takeru Amaya, Ryuichiro Higashinaka and Yugo Takeuchi, "Cognitive Architecture Toward Common Ground Sharing Among Humans and Generative AIs: Trial Modeling on Model-Model Interaction in Tangram Naming Task", in Proceedings of Proceedings of the AAAI Symposium Series (AAAI Symposium 2024), 2024.
  4. Ryo Masumura, Akihiko Takashima, Satoshi Suzuki and Shota Orihashi, "Born-Again Multi-Task Self-Training for Multi-Task Facial Emotion Recognition", in Proceedings of International Conference on Pattern Recognition (ICPR 2024), December 2024.
  5. Naotaka Kawata, Shota Orihashi; Satoshi Suzuki, Tomohiro Tanaka, Mana Ihori and Naoki Maikishima, "Block Refinement Learning for Improving Early Exit in Autoregressive ASR", in Proceedings of Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2024), December 2024.
  6. Kosuke Nishida, Kyosuke Nishida and Kuniko Saito, "Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes", in Proceedings of Conference on Empirical Methods in Natural Language aProcessing (EMNLP 2024), November 2024.
  7. Satoshi Suzuki, Shotaro Tora and Ryo Masumura, "Scene Generalized Multi-View Pedestrian Detection with Rotation-Based Augmentation and Regularization", in Proceedings of International Conference on Image Processing (ICIP 2024), October 2024.
  8. Taiga Yamane, Satoshi Suzuki, Ryo Masumura and Shotaro Tora, "MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with Transformer", in Proceedings of International Conference on Image Processing (ICIP 2024), October 2024.
  9. Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Satoshi Suzuki, Nobukatsu Hojo, Takafumi Moriya and Atsushi Ando, "Unified Multi-Talker ASR Modeling of with and without Target-speaker Enrollment", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  10. Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando and Ryo Masumura, "SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  11. Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno and Ryo Masumura, "Participant-Pair-Wise Bottleneck Transformer for Engagement Estimation from Video Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  12. Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa and Ryo Masumura, "Learning from Multiple Annotator Biased Labels in Multimodal Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (ICASSP 2024), September 2024.
  13. Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa and Ryo Masumura, "Learning from Multiple Annotator Biased Labels in Multimodal Conversation", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH 2024), September 2024.
  14. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa and Ryo Masumura, "TALKING FACE GENERATION FOR IMPRESSION CONVERSION CONSIDERING SPEECH SEMANTICS", in Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024), April 2024.
  15. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa and Ryo Masumura, "Talking Face Generation for Impression Conversion Considering Speech Semantic", in Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024), April 2024.
  16. Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito and Jun Suzuki, "InstructDoc: Zero-shot Generalization of Visual Document Understanding with Instructions", in Proceedings of AAAI Conference on Artificial Intelligence (AAAI 2024), February 2024.

論文

  1. 有山 知希, 鈴木 潤, 鈴木 正敏, 田中 涼太, 赤間 怜奈, 西田 京介, "クイズコンペティションの結果分析から見た日本語質問応答の到達点と課題", 自然言語処理, 2024年 31巻1号, March 2024.
  2. Kazuki Sakai, Koh Mitsuda, Yuichiro Yoshikawa, Ryuichiro Higashinaka, Takashi Minato and Hiroshi Ishiguro, "Effects of Demonstrating Consensus Between Robots to Change User's Opinion", International Journal of Social Robotics 16(7), 2024.

招待・チュートリアル講演

  1. 西田京介, “NTT版大規模言語モデル tsuzumi の取組について", DATA + AI Workd Tour Tokyo, 2024年11月.
  2. 西田京介, “NTT版大規模言語モデル tsuzumi の取組について", weights and biases, Fully Connected Tokyo, 2024年10月.
  3. 西田光甫, “⼤規模⾔語モデルとその学習⼿法", 第4回スマート情報技術研究センターシンポジウム, 2024年10月.
  4. 田中涼太, "大規模言語モデルによる視覚・言語の融合", 第12回岡山大学AI研究会, 2024年7月.
  5. 西田京介, “NTT版LLM「tsuzumi」について", 生成AIフォーラム2024, 2024年6月.
  6. 壹岐太一, “大規模言語モデルに視覚を持たせる : 技術動向とtsuzumiにおける視覚読解", 画像電子学会セミナー, 2024年6月.
  7. 田中涼太, "大規模言語モデルによる文書画像理解の最新動向", LLM-jp勉強会, 2024年6月.
  8. 品川 政太朗,増村 亮,松嶋 達也,"マルチモダリティ革命ー大規模事前学習済みモデルの新たな視点を探るー",人工知能学会全国大会,2024年5月.
  9. 増村 亮,"クロスモーダル表現学習の研究動向: 音声関連を中心として", 日本音響学会春季研究発表会, 2024年3月.
  10. 西田京介, “LLMと共に知の泉を汲んで世に恵みを提供する", 言語処理学会 ワークショップ, 2024年3月.
  11. 西田京介, “NTT版LLM「tsuzumi」の取組について", DEIM2024, 2024年2月.
  12. 西田京介, “大規模言語モデルの基礎・最新動向", 信学会 関西支部イブニングセミナー, 2024年2月.
  13. 西田京介, “NTT版大規模言語モデル「tsuzumi」の取組について", DSAIシンポジウム, 2024年2月.
  14. 西田京介, “LLMのチューニング技術の最新動向", 情処学会 音声言語情報処理研究会, 2024年1月.
  15. 西田京介, “大規模言語モデルの基礎・最新動向", 最強データベース講義, 2024年1月.

表彰

  1. 言語処理学会第30回年次大会(NLP2024) 優秀賞, 田中涼太, 壹岐太一, 西田京介, 齋藤邦子, 鈴木潤, "InstructDoc: 自然言語に基づく視覚的文書理解"
  2. 言語処理学会第30回年次大会(NLP2024) 委員特別賞, 塩野大輝, 宮脇峻平, 田中涼太, 鈴木潤, "大規模視覚言語モデルに関する指示追従能力の検証"
  3. IEEE International Conference on Image Processing 2024 (ICIP2024) Best Industry Paper Award, Taiga Yamane, Satoshi Suzuki, Ryo Masumura and Shotaro Tora, "MVAFormer: RGB-Based Multi-View Spatio-Temporal Action Recognition with Transformer"

2023

国際会議

  1. Keita Suzuki, Satoshi Suzuki, Ryo Masumura, Atsushi Ando and Naoki Makishima, "Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper Body", in Proceedings of the 5th ACM International Conference on Multimedia in Asia (MMAsia 2023), December 2023.
  2. Taku Hasegawa, Kyosuke Nishida, Koki Maeda and Kuniko Saito, "DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning", in Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), December 2023.
  3. Mihiro Uchida, Shota Orihashi, Akihiko Takashima, Yoshihiro Yamazaki and Ryo Masumura, "Open-Set Recognition for Facial-Expression Recognition", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  4. Satoshi Suzuki, Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando and Ryo Masumura, "ONDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training Framework", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  5. Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima and Ryo Masumura, "Distilling Knowledge of Bidirectional Language Model for Scene Text Recognition", in Proceedings of the 23rd International Conference on Image Processing (ICIP 2023), October 2023.
  6. Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando and Ryo Masumura, "Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuray-Robustness Tradeoff", in Proceedings of International Conference on Computer Vision (ICCV 2023), October 2023.
  7. Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka and Shota Orihashi, "Text-to-Text Pre-Training with Paraphrasing for Improving Transformer-based Image Captioning", in Proceedings of the 31st European Signal Processing Conference (EUSIPCO 2023), September 2023.
  8. Mana Ihori, Hiroshi Sato, Tomohiro Tanaka and Ryo Masumura, "Retrieval, Masking, and Generation: Feedback Comment Generation using Masked Comment Examples", in Proceedings of the 16th International Conference on Natural Language Generation (INLG 2023): Generation Challenge, September 2023.
  9. Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura, Saki Mizuno and Nobukatsu Hojo, "Transcribing Speech as Spoken and Written Dual Text Using an Autoregressive Model", in Proceedings of Annual Conference of the International Speech Communication Association (INTERSPEECH), August 2023.
  10. Naoki Makishima, Keita Suzuki, Satoshi Suzuki, Atsushi Ando and Ryo Masumura, "Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  11. Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo and Atsushi Ando, "End-to-End Joint Target and Non-Target Speakers ASR", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  12. Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Hiroshi Sato and Tomohiro Tanaka, "Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer", in Proceedings of the 24th International Speech Communication Association (INTERSPEECH 2023), August 2023.
  13. Saki Mizuno, Nobukatsu Hojo, Satoshi Kobashikawa and Ryo Masumura, "Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation", in Proceedings of the 48th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023), June 2023.
  14. Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato, Taiga Yamane, Takanori Ashihara, Kohei Matsuura and Takafumi Moriya, "Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning", in Proceedings of the 48th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023), June 2023.
  15. Kosuke Nishida, Naoki Yoshinaga (Tokyo Univ.) and Kyosuke Nishida, "Self-Adaptive Named Entity Recognition by Retrieving Unstructured Knowledge", in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2023), May 2023.
  16. Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, Kuniko Saito, "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images", in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI 2023), February 2023.

招待・チュートリアル講演

  1. 西田光甫, "大規模言語モデルとVision-and-Language", 情報論的学習理論ワークショップ, 2023年11月.
  2. 西田京介, 西田光甫, 風戸広史, “大規模言語モデル入門", ソフトウェアエンジニアリングシンポジウム, 2023年8月.
  3. 西田京介, “GPT-4と Vision-and-Languageの未来", 第29回画像センシングシンポジウム, 2023年6月.
  4. 西田京介, 壹岐太一, “Collaborative AI: 視覚・言語・行動の融合”, 第13回 Language and Robotics研究会, 2023年5月.(スライドあり)
  5. 西田京介, 西田光甫, 田中涼太, 斉藤いつみ, “NLPとVision-and-Languageの基礎・最新動向”, 第15回データ工学と情報マネジメントに関するフォーラム チュートリアル, 2023年3月.
  6. 西田京介, “自然言語処理とVision-and-Languageの最新動向”, 東北大学主催 第9回 医学AIセミナー 特別レクチャー, 2023年2月.

表彰

  1. 第256回 情報処理学会自然言語処理研究会 若手奨励賞, 西田光甫, 吉永直樹, 西田京介, "非構造知識検索を用いた自己適応型固有表現認識"
  2. 言語処理学会第29回年次大会(NLP2023) 若手奨励賞, 岡佑依, 田中貴秋, 平尾努, 永田昌明, "球体表面を利用した位置符号化"
  3. 言語処理学会第29回年次大会(NLP2023) 優秀賞:田中涼太, 西田京介, 西田光甫, 長谷川拓, 斉藤いつみ, 齋藤邦子, “SlideVQA: 複数の文書画像に対する質問応答”
  4. 言語処理学会第29回年次大会(NLP2023) 言語資源賞:田中涼太, 西田京介, 西田光甫, 長谷川拓, 斉藤いつみ, 齋藤邦子, “SlideVQA: 複数の文書画像に対する質問応答”
  5. 言語処理学会第29回年次大会(NLP2023) 委員特別賞:西田光甫, 西田京介, 斉藤いつみ, 齋藤邦子, “InstructSum: 自然言語の指示による要約の生成制御”
  6.                              
  7. 言語処理学会第29回年次大会(NLP2023) 委員特別賞 (equal contribution):西田京介, 長谷川拓, 前田航希, 齋藤邦子, “DueT: 視覚・言語のDual-adapter Tuningによる基盤モデル”

2022

国際会議

  1. Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida and Sen Yoshida, "Japanese ASR-Robust Pre-trained Language Model with Pseudo-Error Sentences Generated by Grapheme-Phoneme Conversion", in Proceedings of the 2022 Conference of the International Speech Communication Association (INTERSPEECH 2022), pp. 2688-2692, September 2022.
  2. Fumio Nihei, Ryo Ishii, Yukiko Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama and Takao Nakamura, "Dialogue Acts Aided Important Utterance Detection Based on multiparty and multimodal information", in Proceedings of the 2022 Conference of the International Speech Communication Association (INTERSPEECH 2022), pp. 1086-1090, September 2022.
  3. Kosuke Nishida, Kyosuke Nishida, Shuichi Nishioka, "Improving Few-Shot Image Classification Using Machine- and User-Generated Natural Language Descriptions", in Proceedings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2022) (findings), pp. 1421-1430, July 2022. [arxiv]
  4. Shumpei Miyawaki (Tohoku Univ.), Taku Hasegawa, Kyosuke Nishida, Takuma Kato (Tohoku Univ.), Jun Suzuki (Tohoku Univ.), "Scene-Text Aware Image and Text Retrieval with Dual-Encoder", in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (ACL SRW 2022), pp. 422-433, May 2022.

招待・チュートリアル講演

  1. 西田京介, “視覚と言語に基づく文書理解”, 第42回医療情報学連合大会, 2022年11月.
  2. 西田京介, “深層学習による自然言語処理技術の最新動向とビジネスへの利用例”, 東京大学 総合分析情報学特論, 2022年10月.
  3. 西田京介, “自然言語処理とVision-and-Language”, 2022年度 人工知能学会全国大会(第36回)チュートリアル, 2022年6月.
  4.                              
  5. 田中涼太, "文書画像に対する質問応答技術の最新動向」第2回 AI王 〜クイズAI日本一決定戦〜 , 2022年2月.

表彰

  1. 言語処理学会第28回年次大会(NLP2022) 優秀賞:西田光甫, 西田京介, 西岡秀一, "機械・人の双方が言語で概念を説明可能なFew-shot 画像分類"
  2. 言語処理学会第28回年次大会(NLP2022) 若手奨励賞:田中涼太 "テキストと視覚的に表現された情報の融合理解に基づくインフォグラフィック質問応答"

2021

論文

  1. 鈴木潤(東北大), 松田耕史(東北大), 鈴木正敏(東北大), 加藤拓真(東北大), 宮脇峻平(東北大), 西田京介, "ライブコンペティション:「AI 王~クイズ AI 日本一決定戦~」", 自然言語処理, 2021 年 28 巻 3 号 p. 888-894, September 2021.
  2. 西田京介, "身近になった対話システム:2.機械読解による自然言語理解", 情報処理, 62(10), e7-e11, September 2021.

国際会議

  1. Kosuke Nishida, Kyosuke Nishida, Sen Yoshida, "Task-adaptive Pre-training of Language Models with Word Embedding Regularization", in Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021) (findings), pp. 4546-4553, August 2021. [arxiv]
  2. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Sen Yoshida: "Towards Interpretable and Reliable Reading Comprehension: A Pipeline Model with Unanswerability Prediction." in Proceddings of the 2021 International Joint Conference on Neural Networks ( IJCNN 2021), pp. 1-8, July 2021. [arxiv]
  3. Ryota Tanaka(*), Kyosuke Nishida(*), Sen Yoshida, "VisualMRC: Machine Reading Comprehension on Document Images", in Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI 2021), pp. 13878-13888, a Virtual Conference, February 2021. (*: equal contribution) (full paper, 1696/7911=21.4%) [arxiv] [project]

招待・チュートリアル講演

  1. 西田京介, “深層学習による自然言語処理技術の最新動向とビジネスへの利用例”, 東京大学 総合分析情報学特論, 2021年10月.
  2. 西田京介, “視覚と言語の統合的理解に基づく文書理解と質問応答”, 第48回産業総合研究所人工知能セミナー, 2021年9月.
  3. 西田京介, “人とAIの共生に向けた視覚と言語の融合理解”, NVIDIA AI DAYS, 2021年6月.
  4.                              
  5. 西田京介, “言語と視覚に基づく質問応答の最新動向”, 言語処理学会第27回年次大会ワークショップ AI王 〜クイズAI日本一決定戦〜, 2021年3月.
  6. 西田京介, “自然言語処理とビジョン&ランゲージへの派生”, 日本ロボット学会 ロボット学会セミナー第132回, 2021年2月.

表彰

  1. ICDAR 2021 Competition on Document Visual Question Answering, Task 3 Infographics VQA, runners-up:Ryota Tanaka and Kyosuke Nishida
  2. 2021年度人工知能学会全国大会(JSAI2021) 優秀賞:加来宗一郎, 西田京介, 吉田仙, "BERTにおけるWeightとActivationの3値化の検討"
  3. 言語処理学会第27回年次大会(NLP2021) 最優秀賞 (equal contribution):田中涼太(*), 西田京介(*), 吉田仙, "VisualMRC: 文書画像に対する機械読解"

2020

論文

  1. 大塚淳史, 西田京介, 斉藤いつみ, 西田光甫, 浅野久子, 富田準二 ,"問い返し可能な質問応答:読解と質問生成の同時学習モデル", 日本データベース学会和文論文誌, Vol. 18-J, No. 16, March 2020.

国際会議

  1. Diana Galvan-Sosa (Tohoku Univ.), Jun Suzuki (Tohoku Univ.), Kyosuke Nishida, Koji Matsuda (Tohoku Univ.) and Kentaro Inui (Tohoku Univ.): Seeing the world through text: Evaluating image descriptions for commonsense reasoning in machine reading comprehension, in Proceedings of the Second Workshop on Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN 2020; in conjunction with COLING 2020), December 2020.
  2. Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda and Shoichiro Saito, "A Transformer-based Audio Captioning Model with Keyword Estimation", in Proceedings of the 21st Annual Conference of the International Speech Communication Association (INTERSPEECH 2020), October 2020. [arxiv]
  3. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Hisako Asano and Junji Tomita, "Unsupervised Domain Adaptation of Language Models for Reading Comprehension", in Proceedings of the 12th International Conference on Language Resources and Evaluation (LREC 2020), pp. 5392-5399, May 2020. [arxiv]

招待・チュートリアル講演

  1. 西田京介, “深層学習による自然言語処理技術の最新動向とビジネスへの利用例”, 東京大学 総合分析情報学特論, 2020年10月.
  2. 斉藤いつみ, "ニューラルネットワークを用いた自然言語処理の応用:要約・機械読解モデルの紹介", 2020年10月
  3. 西田京介, “事前学習済言語モデルの動向と展望”, 産総研・東工大実社会ビッグデータ活用オープンイノベーションラボラトリ, 2020年2月.

表彰

  1. 2020年度人工知能学会全国大会(JSAI2020) 優秀賞:斉藤いつみ, 西田京介, 西田光甫, 大塚淳史, 浅野久子, 富田準二, 進藤裕之, 松本裕治, "出力長制御と重要箇所の特定を同時に行う生成型要約"
  2. 2020年度人工知能学会全国大会(JSAI2020) 優秀賞:長谷川拓, 西田京介, 加来宗一郎, 富田準二, "高速な情報検索に向けた文脈考慮型スパース文書ベクトルの獲得"
  3. 言語処理学会第26回年次大会(NLP2020) 若手奨励賞:西田光甫, 西田京介, 斉藤いつみ, 浅野久子, 富田準二, "回答の根拠を解釈可能な機械読解"
  4. 言語処理学会第26回年次大会(NLP2020) 優秀賞:Diana Galvan-Sosa(東北大), 西田京介, 松田耕史(東北大), 鈴木潤(東北大), 乾健太郎(東北大), "テキストを通して世界を見る:機械読解における常識的推論のための画像説明文の評価"

2019

論文

  1. 大塚淳史, 西田京介, 斉藤いつみ, 浅野久子, 富田準二, 佐藤哲司 ,"質問意図の明確化に着目した機械読解による質問応答手法の提案", 人工知能学会論文誌, Vol. 34, No. 5, p. A-J14_1-12, September 2019.
  2. 大塚淳史, 西田京介, 斉藤いつみ, 浅野久子, 富田準二, "質問の意図を特定するニューラル質問生成モデル", 日本データベース学会和文論文誌, Vol. 17-J, No. 6, March 2019. 2018年度日本データベース学会論文賞

国際会議

  1. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, and Takanobu Oba, "Generalized Large-Context Language Models based on Forward-Backward Hierarchical Encoder-Decoder Models", in Proceedings of the 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2019), pp.554-561, December 2019.
  2. Kyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda (Tokyo Univ.), Atsushi Otsuka, Hisako Asano and Junji Tomita, "Multi-style Generative Reading Comprehension", in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pp.2273-2284, July 2019. [arXiv] (long paper, 660/2905=22.7%)
  3. Kosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Itsumi Saito, Atushi Otuka, Hisako Asano and Junji Tomita, "Answering while Summarizing: Multi-task Learning for Multi-hop QA with Evidence Extraction", in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pp.2273-2284, 2335-2345, July 2019. [arXiv] (long paper, 660/2905=22.7%)
  4. Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida, Hisako Asano, and Junji Tomita, "A Simple but Effective Method to Incorporate Multi-turn Context with BERT for Conversational Machine Comprehension", in Proceedings of 1st Workshop on NLP for Conversational AI (NLP4ConvAI 2019; in conjunction with ACL 2019), Florence, Italy, July 2019. [arXiv]
  5. Atsushi Otsuka, Kyosuke Nishida, Itsumi Saito, Hisako Asano and Junji Tomita, "Specific Question Generation for Reading Comprehension", in Proceedings of the AAAI 2019 Reasoning for Complex QA (RCQA) Workshop, Honolulu, Hawaii, USA, January 2019.

招待・チュートリアル講演

  1. 西田京介, “機械読解と自然言語理解”, お茶の水女子大学 理学総論, 2019年11月.
  2. 西田京介, “ACL’19参加報告と事前学習言語モデルの動向 “, xpaper.challenge, 2019年11月.
  3. 西田京介, “深層学習による自然言語処理技術の最新動向とビジネスへの利用例”, 東京大学 総合分析情報学特論, 2019年10月.
  4. 西田京介, “自然言語生成による機械読解”, WebDB Forum 2019 先端研究解説セッション, 2019年9月.
  5. 西田京介, “機械読解の現状と展望”, 言語処理学会第25回年次大会 チュートリアル, 2019年3月.
  6. 西田京介, “機械読解技術の最新動向と実用化へ向けた展望”, 東北大学乾・鈴木研究室 みちのく情報伝達学セミナー, 2019年1月.

表彰

  1. 2018年度日本データベース学会論文賞 ;大塚淳史, 西田京介, 斉藤いつみ, 浅野久子, 富田準二, "質問の意図を特定するニューラル質問生成モデル"
  2. 言語処理学会第25回年次大会(NLP2019) 優秀賞: 西田京介, 斉藤いつみ, 西田光甫, 篠田一聡, 大塚淳史, 浅野久子, 富田準二, "回答スタイルを制御可能な生成型機械読解"
  3. 言語処理学会第25回年次大会(NLP2019) 最優秀ポスター賞:斉藤いつみ, 西田京介, 大塚淳史, 西田光甫, 浅野久子, 富田準二, "クエリ・出力長を考慮可能な文書要約モデル
  4. 第11回データ工学と情報マネジメントに関するフォーラム(DEIM2019) 優秀インタラクティブ賞 :大塚淳史, 西田京介, 斉藤いつみ, 西田光甫, 浅野久子, 富田準二, "問い返し可能な質問応答:読解と質問生成の同時学習モデル"