Dynamic Visual Evidence Reliability Modeling for Multi-Image Multimodal Named Entity Recognition

Abstract

Multi-image multimodal named entity recognition (MNER) not only requires the exploitation of richer visual context, but also faces a more fundamental challenge: different images contribute unequally to entity recognition. Existing multi-image MNER methods typically either aggregate multiple images directly or assume that different images and modalities make approximately stable contributions during fusion. Such static treatment can easily amplify noise in the presence of weakly relevant or redundant images. To address this issue, this paper reconceptualizes multi-image MNER as a task of dynamic visual evidence reliability modeling and proposes a visual evidence reliability modeling framework. Specifically, the framework first encodes multiple images at the image level and constructs a joint visual representation through an inter-image interaction module. It then estimates the visual evidence reliability score based on the consistency between textual and visual representations. Finally, a regulation module consisting of multiple progressive reliability regulation units is introduced to calibrate the contribution of the visual branch in a layer-wise manner, thereby following a simple principle: first assess whether the visual evidence is reliable, and then determine the extent to which it should participate in fusion. Unlike one-shot static fusion, the proposed method can continuously adjust the influence of visual information throughout multi-layer cross-modal interactions, making it better suited to multi-image scenarios where visual support is inherently unstable. Experimental results show that the proposed method achieves competitive performance on the multi-image MNER benchmark. By re-examining the fusion process in multi-image MNER from the perspective of visual evidence reliability, this paper offers a more direct modeling paradigm for multi-image multimodal sequence labeling.

Keywords

References

  1. Devlin J, Chang M W, Lee K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019: 4171-4186.
  2. Luo L, Yang Z, Yang P, et al. An attention-based BiLSTM-CRF approach to document-level chemical named entity recognition[J]. Bioinformatics, 2018, 34(8): 1381-1388.
  3. Xu D, Chen W, Peng W, et al. Large language models for generative information extraction: A survey[J]. Frontiers of Computer Science, 2024, 18(6): 186357.
  4. Zhang Q, Fu J, Liu X, et al. Adaptive co-attention network for named entity recognition in tweets[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2018, 32(1).
  5. Yu J, Jiang J, Yang L, et al. Improving multimodal named entity recognition via entity span detection with unified multimodal transformer[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020: 3342-3352.
  6. Lu D, Neves L, Carvalho V, et al. Visual attention model for name tagging in multimodal social media[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018: 1990-1999.
  7. Sennrich R, Haddow B, Birch A. Neural machine translation of rare words with subword units[C]//Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016: 1715-1725.
  8. Chen D, Li Z, Gu B, et al. Multimodal named entity recognition with image attributes and image knowledge[C]//International Conference on Database Systems for Advanced Applications. 2021: 186-201.
  9. Zhang D, Wei S, Li S, et al. Multi-modal graph fusion for named entity recognition with targeted visual guidance[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(16): 14347-14355.
  10. Wu Z, Zheng C, Cai Y, et al. Multimodal representation with embedded visual guiding objects for named entity recognition in social media posts[C]//Proceedings of the 28th ACM International Conference on Multimedia. 2020: 1038-1046.