Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

FDAM: Frequency-Dynamic Attention Modulation

[ICCV 2025] Official implementation of Frequency-Dynamic Attention Modulation for Dense Prediction(paper link).

FDAM revitalizes Vision Transformers by tackling frequency vanishing. It dynamically modulates the frequency response of attention layers, enabling the model to preserve critical details and textures for superior dense prediction performance.

📰 News

🚀 Key Features

  • Attention Inversion (AttInv): Inspired by circuit theory, AttInv inverts the inherent low-pass filter of the attention mechanism to generate a complementary high-pass filter, enabling a full-spectrum representation.
  • Frequency Dynamic Scaling (FreqScale): Adaptively re-weights and amplifies different frequency bands in feature maps, providing fine-grained control to enhance crucial details like edges and textures.
  • Prevents Representation Collapse: Effectively mitigates frequency vanishing and rank collapse in deep ViTs, leading to more diverse and discriminative features.
  • Plug-and-Play & Efficient: Seamlessly integrates into existing ViT architectures (like DeiT, SegFormer, MaskDINO) with minimal computational overhead.

image-20250214164031767

📈 Performance Highlights

Semantic Segmentation (ADE20K val set)

FDAM boosts performance on various backbones, including CNN-based, ViT-based, and even recent Mamba-based models.

BackboneBase mIoU (SS)+ FDAM mIoU (SS)Improvement
SegFormer-B037.439.8+2.4
DeiT-S42.944.3+1.4
DeiT-III-B (config)51.8 (paper)52.6(model)+0.8

Object Detection & Instance Segmentation (COCO val2017)

Integrated into the state-of-the-art Mask DINO framework, FDAM achieves notable gains with minimal overhead. * indicates reproduced results.

TaskMethodMetricBaseline+ FDAMImprovement
Object DetectionMask DINO (R-50)APbox45.5*47.1+1.6
Instance SegmentationMask DINO (R-50)APmask41.2*42.6+1.4
Panoptic SegmentationMask DINO (R-50)PQ48.7*49.6+0.9

Remote Sensing Object Detection (DOTA-v1.0)

FDAM achieves state-of-the-art results in single-scale settings, showcasing its effectiveness in specialized domains.

BackboneBase mAP+ FDAM mAPImprovement
LSKNet-S77.4978.61+1.12

⚡ Quick Start: Plug-and-Play FDAM Integration

Integrating FDAM into your existing Vision Transformer is straightforward. The core idea is to replace the standard Attention with our AttentionwithAttInv and insert a GroupDynamicScale module after both the attention and MLP blocks. This enhances the model's ability to process frequency information with minimal code changes.

Below is a side-by-side comparison of a standard Transformer Block and our FDAM-enhanced Layer_scale_init_Block.

1. Standard Transformer Block

A typical Block in a Vision Transformer looks like this:

importtorch.nnasnnfromtimm.models.layersimportDropPathfromtimm.models.vision_transformerimportAttention, MlpclassStandardBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm):
super().__init__()
self.norm1=norm_layer(dim)
self.attn=Attention(dim, num_heads=num_heads)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
defforward(self, x):
# Attention blockx=x+self.drop_path(self.attn(self.norm1(x)))
# MLP blockx=x+self.drop_path(self.mlp(self.norm2(x)))
returnx

2. Upgrading to an FDAM Block

To upgrade to FDAM, simply follow these three steps:

  1. Replace Attention with AttentionwithAttInv: This module introduces a high-frequency path.
  2. Add GroupDynamicScale after attention: This module performs frequency scaling on the features from the attention block.
  3. Add GroupDynamicScale after the MLP: This scales the features from the MLP block.

Here is the code for our Layer_scale_init_Block, which encapsulates these changes:

# Assuming AttentionwithAttInv and GroupDynamicScale are defined as in deit_fdam.py# and utility functions nlc_to_nchw, nchw_to_nlc are available.classFdamBlock(nn.Module):
def__init__(self, dim, num_heads, mlp_ratio=4., drop_path=0., norm_layer=nn.LayerNorm, init_values=1e-4):
super().__init__()
# --- Standard components ---self.norm1=norm_layer(dim)
self.drop_path=DropPath(drop_path) ifdrop_path>0.elsenn.Identity()
self.norm2=norm_layer(dim)
self.mlp=Mlp(in_features=dim, hidden_features=int(dim*mlp_ratio))
self.gamma_1=nn.Parameter(init_values*torch.ones(dim))
self.gamma_2=nn.Parameter(init_values*torch.ones(dim))
# --- FDAM-specific upgrades ---# 1. Replace Attention with AttentionwithAttInvself.attn=AttentionwithAttInv(dim, num_heads=num_heads)
# 2. Add GroupDynamicScale after Attention and MLPself.freq_scale_1=GroupDynamicScale(dim=dim)
self.freq_scale_2=GroupDynamicScale(dim=dim)
defforward(self, x, H, W):
# --- Attention block with FDAM ---x_att=self.attn(self.norm1(x))
# Reshape for GroupDynamicScale (N, L, C) -> (N, C, H, W)x_att_reshaped=nlc_to_nchw(x_att, (H, W))
x_att_scaled=self.freq_scale_1(x_att_reshaped) +x_att_reshapedx_att=nchw_to_nlc(x_att_scaled)
x=x+self.drop_path(self.gamma_1*x_att)
# --- MLP block with FDAM ---x_mlp=self.mlp(self.norm2(x))
# Reshape for GroupDynamicScalex_mlp_reshaped=nlc_to_nchw(x_mlp, (H, W))
x_mlp_scaled=self.freq_scale_2(x_mlp_reshaped) +x_mlp_reshapedx_mlp=nchw_to_nlc(x_mlp_scaled)
x=x+self.drop_path(self.gamma_2*x_mlp)
returnx

By replacing your standard Block with this FDAM-enhanced FdamBlock (or our provided Layer_scale_init_Block), you can seamlessly integrate frequency-dynamic modulation into your Vision Transformer models.

🛠 Installation

Our implementation is built on top of MMSegmentation and MMDetection. To ensure compatibility, we recommend following this step-by-step guide.

1. Prerequisites

  • Python 3.8 or higher
  • PyTorch 1.11.0
  • CUDA 11.3

2. Set Up a Virtual Environment (Recommended)

It is highly recommended to use a virtual environment to avoid package conflicts.

# Using conda
conda create -n fdam python=3.8 -y
conda activate fdam
# Using venv
python -m venv venv
source venv/bin/activate

3. Install PyTorch

Install PyTorch and its companion libraries, ensuring they are compatible with CUDA 11.3.

pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 -f https://download.pytorch.org/whl/cu113

4. Install OpenMMLab Dependencies

Install mmcv-full, which is a prerequisite for MMDetection and MMSegmentation.

pip install mmcv-full==1.5.3 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.11.0/index.html

Next, install MMDetection and MMSegmentation.

pip install mmsegmentation==0.25.0

5. Clone and Install FDAM

Finally, clone this repository to your local machine.

git clone https://github.com/your-repo/FDAM.git
cd FDAM

You should now have a complete environment to run the training and evaluation scripts. For more detailed guidance on the MMLab environment, please refer to the official MMSegmentation installation guide.

📖 Citation

If you find this work useful for your research, please consider citing our paper:

Generated code

@InProceedings{chenlinwei2025ICCV,
title={Frequency-Dynamic Attention Modulation for Dense Prediction},
author={Chen, Linwei and Gu, Lin and Fu, Ying},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}

Acknowledgment

This project is built upon the great work from several open-source libraries, including:

We thank their authors for making their code publicly available.

Contact

If you have any questions or suggestions, please feel free to open an issue or contact us at [charleschen2013@163.com]. We welcome any feedback and discussion.

About

ICCV 2025: Frequency-Dynamic Attention Modulation for Dense Prediction

Resources

Stars

85 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages