Summary
The paper jointly learns food categories and ingredients with an Attention Fusion Network that emphasizes discriminative regions and a Food-Ingredient Joint Learning module using balance focal loss to address ingredient imbalance.
Research question
Global dish appearance alone can miss regional ingredient evidence in unstructured Chinese food images, while ingredient labels are imbalanced.
What the paper contributes
- Introduces an Attention Fusion Network (AFN) for region-sensitive food and ingredient features.
- Jointly optimizes fine-grained food-category and ingredient recognition.
- Uses balance focal loss to reduce the effect of ingredient-class imbalance.
Evidence and evaluation scope
Experiments report state-of-the-art ingredient-recognition performance on VIREO Food-172 at publication time. Exact metrics, splits, and comparisons should be cited from the IEEE article.
Scope and limitations
The reported dataset and food domain are specific; ingredient taxonomies, regional cuisine variation, and label imbalance may differ in other settings. The abstract supports AFN, joint learning, and balance focal loss—not only a generic multi-task formulation.
Positioning for related work
This work connects fine-grained food recognition with explicit ingredient prediction and regional attention, treating ingredient composition as both an auxiliary semantic signal and a target affected by imbalance.
Related-work context
Liu et al. jointly model fine-grained food categories and ingredients using an Attention Fusion Network and balance focal loss to capture discriminative regions and mitigate ingredient imbalance.
A concise, neutral description of how this paper can be situated in related work.
Official abstract
Fine-grained food recognition is the detailed classification that provides more specialized and professional attribute information of food. It is the basic work to realize healthy diet recommendations and cooking instructions, nutrition intake management, and cafeteria self-checkout system. Chinese food lacks structured information, and ingredients composition is an important consideration. The current approaches mostly focus on global dish appearance without any analysis of ingredient composition and fully considering the attention of regional features. In this paper, we propose an Attention Fusion Network (AFN) and Food-Ingredient Joint Learning module for fine-grained food and ingredients recognition. The AFN first focuses on the food discrimination region against unstructured defeat and generates the feature embeddings jointly aware of the ingredients and food. The Food-Ingredient Joint Learning module aims at alleviating the issue of ingredients imbalance. Therefore, we propose a balance focal loss to optimize the feature expression ability of the network for ingredients. In experiments, the results of ingredients recognition show the state-of-the-art performances on fine-grained Chinese food dataset VIREO Food-172.
The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.
Evidence references
| What to verify | Location in the paper |
|---|---|
| Problem statement | Abstract |
| Method and contributions | Abstract; Attention Fusion Network, joint-learning module, and balance focal loss |
| Evaluation statement | Abstract; VIREO Food-172 experiments |
Primary source: IEEE version of record (IEEE early access 2020-08-28; TCSVT 31(6), June 2021).
How to cite
Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.
Chengxu Liu, Yuanzhi Liang, Yao Xue, Xueming Qian, and Jianlong Fu. “Food and Ingredient Joint Learning for Fine-Grained Recognition.” IEEE Transactions on Circuits and Systems for Video Technology (2021), 31(6), 2480-2493. https://doi.org/10.1109/TCSVT.2020.3020079.
Reuse policy
Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.