论文解读 · 作者已确认

Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification

作者: , Linchao Zhu, Xiaohan Wang, Yi Yang

IEEE Transactions on Neural Networks and Learning Systems · 2024 · 卷 35(5) · 页 7048-7059

fine-grained visual classificationhard examplesloss modulationgeneralizationoverfitting

论文信息

作者
Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang
推荐论文引用
Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang. “Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification.” IEEE Transactions on Neural Networks and Learning Systems (2024), 35(5), 7048-7059. https://doi.org/10.1109/TNNLS.2022.3213563.

版本日期

Online 首发
2022-11-21
来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

MHEM 的全称是 Moderate Hard Example Modulation:它为调制损失定义条件,强调有信息量的困难样本,同时避免极端样本的影响不断增大并被细粒度分类器直接记忆。

研究问题

细粒度网络可能完全拟合训练集中的极难样本,却仍无法识别测试集中的相似难例,说明模型学到的是记忆而非可迁移判别能力。

论文贡献

  • 提出三个条件并定义适度处理 hard example 的通用 modulated loss 形式。
  • 将该形式实例化为简单而强的 FGVC baseline。
  • 说明这一 baseline 可接入现有细粒度识别方法。

证据与评测范围

论文在 CUB-200-2011、Stanford Cars 与 FGVC-Aircraft 上报告一致改进;精确比较应引用 TNNLS 最终版本和实验表格。

适用范围与局限

MHEM 不会自动判断每个极端样本究竟是错标、异常还是确有信息,而是通过 loss design 控制其影响;证据主要集中在细粒度分类。

Related work 定位

MHEM 是 FGVC 的 loss modulation 与泛化 baseline。缩写指 Moderate Hard Example Modulation,不是 Mining;论文 2022 年 online,最终刊于 2024 年 TNNLS 35(5)。

Related Work 表述

Liang 等提出 Moderate Hard Example Modulation(MHEM),通过限制极端困难样本的过度影响,改善细粒度视觉分类的泛化。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

Though significant progress has been achieved on fine-grained visual classification (FGVC), severe overfitting still hinders model generalization. A recent study shows that hard samples in the training set can be easily fitted, but most existing FGVC methods fail to classify some hard examples in the test set. The reason is that the model overfits those hard examples in the training set, but does not learn to generalize to unseen examples in the test set. In this paper, we propose a Moderate Hard Example Modulation (MHEM) strategy to properly modulate the hard examples. MHEM encourages the model to not overfit hard examples and offers better generalization and discrimination. First, we introduce three conditions and formulate a general form of a modulated loss function. Second, we instantiate the loss function and provide a strong baseline for FGVC, where the performance of a naive backbone can be boosted and be comparable with recent methods. Moreover, we demonstrate that our baseline can be readily incorporated into the existing methods and empower these methods to be more discriminative. Equipped with our strong baseline, we achieve consistent improvements on three typical fine-grained visual classification datasets, i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft. We hope the idea of Moderate Hard Example Modulation will inspire future research work toward more effective fine-grained visual recognition.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; Moderate Hard Example Modulation formulation
评测结论Abstract; CUB-200-2011, Stanford Cars, and FGVC-Aircraft experiments

主要核验来源: IEEE version of record (IEEE early access 2022-11-21; TNNLS 35(5), May 2024).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang. “Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification.” IEEE Transactions on Neural Networks and Learning Systems (2024), 35(5), 7048-7059. https://doi.org/10.1109/TNNLS.2022.3213563.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源