Mobile QR Code QR CODE

2025

Reject Ratio

81.5%

Title Cross Deformable Fusion with Image-Aware Text Prompts for Semantic Segmentation
Authors (So-Yeon Jang) ; (Jong-Ok Kim)
DOI https://doi.org/10.5573/IEIESPC.2026.15.4.532
Page pp.532-540
ISSN 2287-5255
Keywords Semantic segmentation; Multi-modal learning; Text-guided segmentation; Discrete wavelet transform; Deformable convolution
Abstract Recent vision-language models are increasingly applied to dense prediction tasks such as semantic segmentation, where textual input provides high-level semantic guidance. A common strategy is to use cross-attention mechanisms to integrate image and text features. However, such methods often suffer from limited spatial adaptability and substantial computational overhead. To resolve these issues, we propose an architecture composed of a Wavelet-Aware Context Construction (WACC) module and a Cross Deformable Fusion (CDF) module. WACC generates enriched textual representations by decomposing visual features into multiple frequency components via Discrete Wavelet Transform (DWT). It captures both global semantic layout and fine-grained structural cues. These informative text features are subsequently fused with image features using CDF, which employs deformable convolution to achieve spatially adaptive and content-aware alignment across modalities. This design enables more precise and efficient cross-modal interaction, leading to improved segmentation performance. Experimental results demonstrate that our method achieves an mIoU of 0.790, exceeding the baseline of 0.777 and confirming the effectiveness of the WACC and CDF modules.