Evidence map›Paper›PMID 38257093›Full record

ArticleNutrients2024

A Lightweight Hybrid Model with Location-Preserving ViT for Efficient Food Recognition.

Guorui Sheng, Weiqing Min, Xiangyi Zhu, Liang Xu, Qingshuo Sun, Yancun Yang, Lili Wang, Shuqiang Jiang

Open access · goldAbstract read
In one paragraph

Article in Nutrients, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
4.1field-weighted citation impact, top 6% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 22 citations in OpenAlex.

  1. Article
  2. Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors at 2 institutions in 1 country.

Guorui ShengSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.ORCID 0000-0001-6790-0239
Weiqing MinKey Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China.ORCID 0000-0001-6668-9208
Xiangyi ZhuSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.
Liang XuSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.
Qingshuo SunSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.
Yancun YangSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.
Lili WangSchool of Information and Electrical Engineering, Ludong University, Yantai 264025, China.
Shuqiang JiangKey Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China.
Ludong University · CNChinese Academy of Sciences · CN

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Food-image recognition plays a pivotal role in intelligent nutrition management, and lightweight recognition methods based on deep learning are crucial for enabling mobile deployment. This capability empowers individuals to effectively manage their daily diet and nutrition using devices such as smartphones. In this study, we propose an Efficient Hybrid Food Recognition Net (EHFR-Net), a novel neural network that integrates Convolutional Neural Networks (CNN) and Vision Transformer (ViT). We find that in the context of food-image recognition tasks, while ViT demonstrates superiority in extracting global information, its approach of disregarding the initial spatial information hampers its efficacy. Therefore, we designed a ViT method termed Location-Preserving Vision Transformer (LP-ViT), which retains positional information during the global information extraction process. To ensure the lightweight nature of the model, we employ an inverted residual block on the CNN side to extract local features. Global and local features are seamlessly integrated by directly summing and concatenating the outputs from the convolutional and ViT structures, resulting in the creation of a unified Hybrid Block (HBlock) in a coherent manner. Moreover, we optimize the hierarchical layout of EHFR-Net to accommodate the unique characteristics of HBlock, effectively reducing the model size. Our extensive experiments on three well-known food image-recognition datasets demonstrate the superiority of our approach. For instance, on the ETHZ Food-101 dataset, our method achieves an outstanding recognition accuracy of 90.7%, which is 3.5% higher than the state-of-the-art ViT-based lightweight network MobileViTv2 (87.2%), which has an equivalent number of parameters and calculations.

Indexed as

Extracellular TrapsFoodCognitionHumansIntelligenceNutritional StatusReceptor Protein-Tyrosine KinasesReceptor Protein-Tyrosine Kinasesfood recognitionglobal featurelightweightnutrition managementViT

Identifiers

PMID38257093
PMCPMC10819383
OpenAlexW4390668715

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.