Paper overview · author-verified

Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification

Authors: , Linchao Zhu, Xiaohan Wang, Yi Yang

IEEE Transactions on Neural Networks and Learning Systems · 2024 · vol. 35(5) · pp. 7048-7059

fine-grained visual classificationhard examplesloss modulationgeneralizationoverfitting

Publication details

Authors
Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang
Recommended paper citation
Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang. “Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification.” IEEE Transactions on Neural Networks and Learning Systems (2024), 35(5), 7048-7059. https://doi.org/10.1109/TNNLS.2022.3213563.

Version dates

First published online
2022-11-21
Source checked
2026-07-31
Verification status
Author-verified on 2026-07-31

Summary

MHEM means Moderate Hard Example Modulation: it defines conditions for a loss that emphasizes informative hard samples without continually increasing the influence of extreme cases that a fine-grained classifier may simply memorize.

Research paths

How this paper contributes to the site's broader research map.

Research question

Fine-grained networks can perfectly fit very hard training examples yet fail on similarly hard test cases, indicating memorization rather than transferable discrimination.

What the paper contributes

  • Defines three conditions and a general form for modulated losses that treat hard examples moderately.
  • Instantiates the formulation as a simple, strong FGVC baseline.
  • Shows that the baseline can be incorporated into existing fine-grained methods.

Evidence and evaluation scope

The paper reports consistent improvements on CUB-200-2011, Stanford Cars, and FGVC-Aircraft. Use the final TNNLS volume and experiment tables for precise comparisons.

Scope and limitations

MHEM does not automatically identify whether every extreme sample is mislabeled, atypical, or genuinely informative; it controls influence through loss design. Its evidence is centered on fine-grained classification.

Positioning for related work

MHEM is a loss-modulation and generalization baseline for FGVC. The acronym expands to Moderate Hard Example Modulation, not 'Mining'; the article appeared online in 2022 and in TNNLS volume 35, issue 5 in 2024.

Related-work context

Liang et al. propose Moderate Hard Example Modulation (MHEM), a loss-modulation strategy that limits overemphasis on extreme hard samples to improve generalization in fine-grained visual classification.

A concise, neutral description of how this paper can be situated in related work.

Official abstract

Though significant progress has been achieved on fine-grained visual classification (FGVC), severe overfitting still hinders model generalization. A recent study shows that hard samples in the training set can be easily fitted, but most existing FGVC methods fail to classify some hard examples in the test set. The reason is that the model overfits those hard examples in the training set, but does not learn to generalize to unseen examples in the test set. In this paper, we propose a Moderate Hard Example Modulation (MHEM) strategy to properly modulate the hard examples. MHEM encourages the model to not overfit hard examples and offers better generalization and discrimination. First, we introduce three conditions and formulate a general form of a modulated loss function. Second, we instantiate the loss function and provide a strong baseline for FGVC, where the performance of a naive backbone can be boosted and be comparable with recent methods. Moreover, we demonstrate that our baseline can be readily incorporated into the existing methods and empower these methods to be more discriminative. Equipped with our strong baseline, we achieve consistent improvements on three typical fine-grained visual classification datasets, i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft. We hope the idea of Moderate Hard Example Modulation will inspire future research work toward more effective fine-grained visual recognition.

The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.

Evidence references

What to verifyLocation in the paper
Problem statementAbstract
Method and contributionsAbstract; Moderate Hard Example Modulation formulation
Evaluation statementAbstract; CUB-200-2011, Stanford Cars, and FGVC-Aircraft experiments

Primary source: IEEE version of record (IEEE early access 2022-11-21; TNNLS 35(5), May 2024).

How to cite

Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.

Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, and Yi Yang. “Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification.” IEEE Transactions on Neural Networks and Learning Systems (2024), 35(5), 7048-7059. https://doi.org/10.1109/TNNLS.2022.3213563.

Reuse policy

Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.

Primary sources and resources