Mobile QR Code QR CODE

2025

Reject Ratio

81.5%


  1. (School of Culture, Media and Art Design, Liaoning University of Technology, Jinzhou 121000, China)
  2. (School of Communication Engineering, Liaoning Railway Vocational and Technical College, Jinzhou 121000, China)



Deep learning, Image recognition, Advertisement, Algorithm optimization, Migration network, Res-Net152

1. Introduction

In recent years, with the rapid progress of artificial intelligence technology and the booming rise of clothing advertising e-commerce, advertising image recognition technology, as an important bridge connecting consumers and commodities, is gradually becoming a key force driving industry change. The rapid development of this field not only greatly enriches the form of advertisement, but also stimulates the public’s high expectations for personalized and precise content delivery. In the field of apparel advertising, image style recognition technology is especially critical, which can not only quickly capture and analyze the fashion elements in advertisements, but also make intelligent recommendations based on user preferences, thus greatly enhancing user experience and brand loyalty. However, in the pursuit of image style recognition accuracy, the field of apparel advertising image recognition also faces many challenges. First of all, clothing advertising images often contain rich details and changing styles, from retro elegance to modern avant-garde, from simple and fresh to complex and gorgeous, each style requires fine classification and recognition. This kind of extremely high demand for image classification fineness makes the traditional image processing and feature extraction methods seem incompetent. Secondly, subtle features in apparel advertising images, such as fabric texture, color matching, pattern design, etc., are often difficult to be effectively captured and distinguished by simple algorithmic models, which further increases the complexity of recognition. In order to break through this technical bottleneck, our team is committed to optimizing and perfecting the style recognition method for clothing advertisement images based on Res-Net 152, which, as a bright pearl in the field of deep learning, has achieved remarkable results in the field of image recognition due to its powerful feature extraction capability and efficient computational efficiency. However, in the face of the special needs of clothing advertisement image style recognition, we realize that relying solely on the infrastructure of Res-Net 152 is still insufficient, and we need to introduce more advanced deep learning algorithms and optimization strategies.

Specifically, we plan to make improvements in the following aspects: first, we will explore and integrate the attention mechanism into the Res-Net 152 model. The attention mechanism can mimic the selective attention ability of the human visual system, which enables the model to focus more on key regions and features in the image during the recognition process, thus improving the accuracy and efficiency of recognition. Secondly, we will adopt a transfer learning approach by utilizing the pre-trained Res-Net 152 model on a large image dataset as a starting point, which will be fine-tuned in a way to make it better adapted to the specific task of apparel advertisement image style recognition. In addition, we will also consider introducing a multimodal learning strategy that combines multiple sources of information, such as text descriptions and user feedback, to further improve the comprehensiveness and accuracy of the recognition.

Through the implementation of these optimization strategies, we expect to build a more intelligent and efficient clothing advertisement image style recognition system. The system will not only be able to accurately identify the style elements in the advertisement images, but also make personalized recommendations based on the user’s historical behavior and real-time preferences, bringing an unprecedented innovative experience to the apparel advertising industry. We believe that as technology continues to advance and application scenarios continue to expand, apparel advertising image style recognition technology will play a more important role in promoting the intelligent development of the industry, and bring consumers a richer, more diverse and personalized shopping experience.

This paper is structured as follows. In Section 2, a literature review of neural networks and image recognition applied to the field of advertising is presented. In Section 3, an improved apparel advertisement recognition model is constructed. In Section 4, the experimental test of the recognition algorithm of this paper is carried out. in Section 5, the summary and analysis are carried out.

2. Literature Review

For example, Chao et al. [1- 3] skillfully used feature descriptors such as Histogram of Oriented Gradients (HOG) and Local Binary Patterns (LBP) to portray the unique styles of apparel advertisement images, and based on the similarity metrics between these features, a personalized recommendation of apparel advertisement styles was achieved. Personalized recommendation of clothing advertisement styles is achieved. This approach demonstrates the important role of feature engineering in image recognition. In addition, there are also researchers such as Zhuang et al. [4], who improved the Canny algorithm by focusing on recognizing and classifying the structural features of clothing advertisement styles, which effectively improves the ability of recognizing the structural styles of advertisement images. Gao et al. [5], on the other hand, took an alternative approach by adopting an optimized HSR-FCN architecture, innovatively integrating the environmental Region Proposal Network (RPN) and HyperNet network into the R-FCN framework. By changing the image feature learning method and introducing a spatial transformation network to spatially transform and align the input image and feature map, the feature learning ability of multi-view and deformed clothing advertisements is enhanced, which effectively solves the problem of recognizing deformed advertisement images.

Recent advancements in image recognition have increasingly relied on deep neural networks (DNNs) due to their ability to automatically learn and extract image features, eliminating the need for manually designed feature descriptors. Ko-Ziarski et al. [7] explored how incorporating different types and severities of noise into DNNs can enhance image recognition performance, highlighting the superiority of this approach over methods that ignore noise. Shubathra et al. [8] compared three methods—MLP, Network Architecture, and ELM—for recognizing clothing advertisement images. Their findings revealed that ELM outperforms the others in terms of speed and accuracy. Building on these concepts, Wang et al. [9] proposed a texture image recognition method using deep convolutional neural networks (CNNs) combined with transfer learning. Their approach involved replacing fully connected layers with global average pooling, optimizing the transfer learning model through fine-tuning, and determining the best combination of tunable and frozen layers to achieve optimal accuracy in recognizing textured apparel images. Similarly, Elleuch et al. [10] employed the Inception-v3 network, using transfer learning to enhance training efficiency and accuracy in recognizing clothing styles from a dataset of 80,000 images. Despite these advancements, current research predominantly focuses on basic feature extraction for coarse classification, which is insufficient for recognizing subtle style differences in clothing advertisement images. To address this, Lin et al. [11] introduced a dual-convolutional network model with parallel pathways for feature extraction, applying dual regularity pooling to capture correlations between features and enhance the recognition of subtle factors. However, this approach significantly increases computational complexity and resource consumption.

This model uses two parallel network pathways to extract image feature degrees and then adopts the dual regularity pooling method to calculate the correlation between these two parallel feature degrees, this process is used to filter the feature degrees of subclasses of subtle factors so that it can fully exploit the subtle factor feature degrees [12- 14]. However, the dual regularity network architecture model employs two parallel network pathways to extract the feature degree, which leads to an exponential increase in the parameter mechanism and computation, and a corresponding increase in the training and optimization time and computational resource consumption [15- 17]. Based on the above existing problems [18- 22], this paper proposes a method for recognizing the style of apparel advertisement images based on the improved Res-Net152 network and migration degree learning. For the Res-Net152 network [23- 25], which has been optimized by early exercise on a collection of Image-Net arrays, the migration degree learning method of sharing model representation parameters is used to migrate the parameters of the model representation of the early exercise optimization into the improved Res-Net152 network.

However, although traditional style recognition methods for apparel advertising images have played a positive role in promoting the intelligentization process of the industry, they have shown their limitations in terms of accuracy and processing efficiency. Specifically, these methods are often difficult to accurately capture all the subtle style features when faced with complex and changing clothing advertisement images, resulting in biased recognition results and difficulty in meeting the demand for high precision. Meanwhile, when dealing with large-scale or real-time demanding image data, traditional methods have high computational complexity and relatively slow processing speed, which is obviously not efficient enough for small-batch array acquisition scenarios that require rapid response to market changes. Therefore, the development of more efficient and accurate clothing advertisement image style recognition technology has become a key issue to be solved in the current industry.The aim of our work is to further optimize and refine the method for recognizing the style of clothing advertisement images based on Res-Net 152 with transfer learning technique. We plan to do so by introducing more advanced deep learning algorithms and optimization strategies. Ultimately, we expect to provide more intelligent and efficient image style recognition solutions for the apparel advertising industry and promote the further intelligent development of the industry!

The contribution of this study is to propose an innovative method based on Res-Net152 with transfer learning. The method enables efficient recognition for girl’s clothing advertisement image styles. By optimizing the network architecture, the recognition accuracy is significantly improved and the training time is drastically shortened to reduce the cost. This technique is not only applicable to diverse advertising scenarios, but also provides strong support for advertising image design and localization, demonstrating the potential for wide application and high efficiency and accuracy in the field of advertising image recognition.

3. Improved Model

Clothing advertisement image style recognition not only needs to consider the basic feature degree information such as the color and texture of the clothing advertisement itself. It also needs to take into account subtle factors such as the style of the clothing advertisement. The rich feature degree information of clothing advertisement style can significantly improve recognition accuracy. To obtain rich feature degree information, can be solved by augmentation network. Compared to networks such as Alex-Net and Google-Net [26- 29], the Res-Net network allows the network organization to learn the pre-discrepancy due to the introduction of the residual group end. It allows it to learn the feature degree information more efficiently. Also, the residual group end structure adds inputs directly to the group outputs through leapfrog connections. The class gradient can be easily propagated back to shallower groups, alleviating the class gradient dissipation problem and also helping to prevent network regression. Thus Res-Net can more easily train to optimize very deep neural networks without problems such as class gradient dissipation and regression. The residual group end structure is shown in Fig. 1.

Fig. 1. Schematic of residual group end structure.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig1.png

According to the process embodied in Fig. 1, the user will have the options of selecting the initial group of images to be recognized, selecting the image to specify the environment area, and selecting the local image. Advertising image category information from different sources can be recognized and processed, not only limited to the camera’s shooting performance, which can optimize the user experience and improve user satisfaction. In this chapter, a generalized definition of video image information group segmentation is given: according to the style requirements of Caffe-x for the input image information group format, the input video image is segmented into sub-environmental regions based on the eigen-degree value channel. The output of the image information group segmentation is the coordinate information of a rectangle-like box of potential locations of possible objects. After segmenting the image based on this coordinate value it is input into the Caffe trained and optimized C class network class divider for recognition. If it is the target object class then the environment region will be identified with the corresponding coordinate information fed back to the source video image.

Due to the difference in the number of network groups, the Res-Net network includes Res-Net18, Res-Net34, Res-Net50, Res-Net101, and Res-Net152, etc., Res-Net18 and Res-Net34 adopt the Ba-14sicBlock structure, and Res-Net50, Res-Net101, and Res-Net152. Net101, and Res-Net152. To deal with the complex clothing advertisement image style recognition task, the style feature degree is extracted to a deeper group of times. In this paper, the Res-Net152 network is chosen to build the model, whose structure is shown in Fig. 2, and the total number of groups is 152, including a 7*7 convolutional group, Bottleneck-structures from stage 1 to stage 4, and fully connected groups with average-pooling and SoftMax functions. Among them, stage 1 to stage 4 has 3 residual group ends, stage 2 has 8 residual group ends, stage 3 has 36 residual group ends, stage 4 has 3 residual group ends, and each residual group end contains 3 convolutional groups. The specific network flow architecture of the processing process is shown in Fig. 2.

Fig. 2. Process network flow architecture.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig2.png

Improving the accuracy of style recognition of clothing advertisement images requires the network to extract the style feature degree of more subtle factors. To achieve this goal, this paper proposes a method to improve the first group class structure of the network. That is, the 7×7 convolution center that extracts the style characteristic degree on the input clothing advertisement image is replaced with three 3×3 convolution center combination groups, keeping the same step amplitude and padding setting in the convolution operation. The network’s first group structure before and after improvement is shown in Fig. 3.

Fig. 3. Before and after improvements to the network headgroup structure.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig3.png

Using a combined group of three 3*3 convolution centers instead of one 77 convolution center can keep the size of the perceptual and output indegree maps in the network constant. The formula is shown in Eq. (1):

(1)
$ F(i) = (F(i+1)-1) * Strid + Ksize, $

where $F(i+1)$ denotes the sensibility of group $i+1$, $F(i)$ denotes the sensibility of group $i$, Stride denotes the step size, and Ksize denotes the size of the convolution center. In the experiment, the initial value of $F(i +1)$ is set to 2, and the initial value of Stride is 1. Group 1 of the network goes through a $7 * 7$ convolution center. According to Eq. (1): $F(1) = (2-1) * 1+7 = 8;$

The 1st group of the network passes through three $3 * 3$ convolutional center combination groups, which is known according to Eq. (1):

$F(1) = (2-1) * 1+3 = 4,$

$F(2) = (4-1) * 1+3 = 6,$

$F(3) = (6-1) * 1+3 = 8.$

Therefore, it can be seen that using a combined group of three $3*3$ convolutional centers instead of one $7*7$ convolutional center can keep the size of the sensibility constant, and the group of three $3 * 3$ convolutional centers from the structure can increase the depth of the network, capturing the style feature degree at multiple scales. Introducing more unconventionality reduces overfitting and improves the performance of the network.

Res-Net residual network consists of multiple residual learning modules superimposed on each other. When the input array of clothing advertisement images enters the Res-Net residual network, it needs to go through a series of processing. Firstly, the convolution group (Conv) extracts the feature degree of the input clothing advertisement images and then enhances the unconventional fitting level of the network through the unconventional starting function group (Relu). The array is processed by the Batch Normalized Scale Group (BN) for normalizing the scale. The result of the processing is then fed into multiple residual modules which further process the array through batch-normalized proportional groups (BN) and multiple fully connected groups. Finally, the output clothing advertisement image is obtained.

In deep-group networks, the original residual modules may encounter problems with class gradient dissipation or class gradient explosion. To solve the potential problem, this paper proposes a method to vary the way of combining the residual modules. That is, the combination of 11convolution group (Conv) + batch-normalized proportional group (BN) + unconventional starting function group (Relu)” is replaced by the “combination of 11batch-normalized proportional group (BN) + unconventional starting function group (Relu) + convolution group (Conv)". The change in combination introduces a pre-start structure, which helps to normalize the start values by placing the BN group at the beginning of the non-conventional branch so that they are within a reasonable range to reduce the class gradient dissipation problem and make the network easier to train for optimization. Same as the above method, we randomly tested the residual degree of improved recognition of sportswear advertisements in a certain environment. The results are shown in Fig. 4, where we find that the improved method has intentional environment region delineation results in both grayscale and color recognition environments. This proves that the improved method has an excellent level of image recognition and processing for apparel advertisements.

Fig. 4. Improved test results for the residual module.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig4.png

To demonstrate the effect of different network improvement methods of Res-Net152 on the style recognition effect of clothing advertisements, the representation parameters of the Res-Net152 network model, which is well-trained and optimized on the Image-Net array collection, are migrated to the improved network. The collection of girls’ clothing advertisement arrays collected in this paper is input into different improved networks for training optimization. The style recognition accuracy of different network improvement methods is obtained as shown in Table 1.

Table 1. Recognition accuracy under different network styles.

Improved methodology Accuracy (%)
Improvement of the first group 89.4
Improvement of the residual module 91.7
Simultaneous improvements 94.2

As can be seen from Table 1, combining the improved network first group structure method and the improved residual module method, this time the network achieves the highest style recognition accuracy of 94.2% for the collection of girl’s clothing advertisement arrays. Therefore, in this paper, the network combining two improved methods is used for clothing advertisement style recognition.

4. Test

4.1. Pre-experimental Treatment

The collection of girls’ clothing advertisement image arrays used in this paper comes from major e-commerce companies or platforms, due to the type variety of girls’ clothing advertisement styles. Through expert advice and querying related information, this paper selects the four most representative types of girls’ clothing advertisement styles, which are Cute-Style, Sports-Style, College-Style, and Ethnic-Style, and each style contains 300 images, totaling 1,200 advertisement images of girls’ clothing. Among them, 80% of the images are used as the training and optimization set, which is used to train and optimize the network architecture model and adjust the model representation parameters; 20% of the images are used as the validation set, which is used to validate the training and optimization effect of the network model and fine-tune the model representation parameters. In labeling the array collection, only the girl’s clothing advertisement images need to be put into the collection of files for the corresponding style classification, and no additional labeling is required. The partial images of the array collection are shown in Figure 5, which presents the diversity of girls’ clothing advertisement styles in the array collection. At the same time, due to the interference of the shooting angle, lighting, folds, background, and other factors of the clothing advertisements in the array collection, the difficulty of recognizing the style of the girls’ clothing advertisement images is increased to a certain extent.

Fig. 5. Breadth of employee valuation - job network structure.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig5.png

4.2. Array Pre-processing and Enhancement

Variations in the array collection in terms of clothing advertisement shooting angle, illumination, etc. introduce noise and variance, which increases the complexity of the network’s task of recognizing the style of girl’s clothing advertisement images. To address this challenge, array pre-disposition and enhancement techniques can be used to normalize the array collection to reduce the impact of this factor.

4.3. Experiments and Methods

The experimental environment is based on Windows 11, AMD R7-6800H processor, 16GB RAM and 512GB SSD, and the PyTorch deep learning framework is used for training and testing. The specific parameters are set as follows: the Batch Size is set to 32, the initial learning rate is set to 0.001, the Adam optimizer is used, and the number of iterations is set to 200. These parameters can continue to be adjusted after the algorithm is optimized. The recognition accuracy and loss degree function obtained after training optimization are visualized by using the Python-Mat library. In this paper, a comparison experiment is set up to train and optimize the Res-Net152 network without mobility learning, the improved Res-Net152 network without mobility learning, the Res-Net152 network with mobility learning, and the improved Res-Net152 network with mobility learning using the training optimization set of girls’ clothing advertisements, respectively. After the completion of training optimization, to verify the impact of the mobility learning method and the improved network on the recognition of the style of girls’ clothing advertisement images, the degree of recognition accuracy (val) and the loss degree function (loss) of the Res-Net152 and the improved Res-Net152 using mobility learning and without mobility learning are recorded once every 5 iterations. The network model learning coefficient (lr) was set to 0.0001, the batch size (batch_size) to 16, and the total number of iterations (epoch) to 200. The version of the advanced exercise optimization model is Res-Net152-394f9c45.pth

To optimize clothing advertisement image style recognition, we employ transfer learning by migrating efficient representation parameters from advanced models to our improved ResNet-152 network. This strategy enhances the network by leveraging successful experiences from different domains, aiming to standardize image acquisition, reduce external interference, and ensure high-quality data input. After preprocessing, standardized images are input into the improved ResNet-152 network for training. Through iterative optimization, the network learns key style features, enabling efficient and accurate recognition in complex image environments. We validated the method through exhaustive algorithmic implementation, prediction accuracy evaluation, and computational efficiency comparisons, demonstrating its effectiveness.

Figs. 6(a) and 6(b) visually demonstrate the significantly improved prediction performance of the improved Res-Net152: its prediction value matches the real value much better, and the error rate in the data test is drastically reduced from the original Res-Net152’s 3.08% to 1.87%, which is a reduction of 1.21 percentage points in error. This result fully demonstrates the effectiveness of combining migration learning with preprocessing enhancement techniques. Further, Fig. 6c illustrates the comparative results in terms of computational efficiency. In several computational efficiency test trials, the improved Res-Net152 demonstrates the highest computational efficiency, which is about one order of magnitude higher compared to the original Res-Net152 and other Res-Net variants. This achievement not only reflects the optimization results of the network structure, but also highlights the great potential of the algorithm in practical applications.

Fig. 6. Improved Res-Net152.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig6.png

5. Results

The training optimization sets of girl’s clothing advertisement images are inputted into the Res-Net152 network without mobility learning, the improved Res-Net152 network without mobility learning, the Res-Net152 network with mobility learning, and the improved Res-Net152 network with mobility learning proposed in this paper, respectively, and are trained and optimized to obtain the corresponding loss degree function (loss). After the training optimization is completed, its validation group is input to the above four networks for style recognition respectively, and the corresponding style recognition accuracy degree (val) is obtained. Figs. 7(a)-7(d) show the style recognition accuracy (val) and loss function (loss) obtained by the four networks on the same set of girl’s clothing advertisement arrays, respectively. Fig. 7(a) shows the loss function and recognition accuracy of the Res-Net152 network without using mobility learning for training and optimization on the collection of girl’s clothing advertisements array, from the curve graphs, it can be seen that when this network is trained and optimized on the collection of small batch of arrays, the accuracy of clothing advertisements styles recognition is low and the loss function is high, and the network model is iterated for a total of 80 times, of which the accuracy is highest when iterating to the 68th iteration, reaching only 65.0% and the accuracy is high. The initial Res-Net152 network is poorly trained and optimized, the loss function and recognition accuracy fluctuate greatly, and the model may have been overfitted. Fig. 7(b) shows the loss function and recognition accuracy of the improved Res-Net152 network without migration learning when trained on a collection of girl’s clothing advertisements. from the curve graph, it can be seen that when this network is trained on a collection of small batch arrays, the accuracy is the highest at the 40th iteration, reaching 85.9%, and the corresponding loss function is 0.258. Compared with the Res-Net152 network without migration learning, the loss function is 0.927, and the loss function is 0.258. Compared to the Res-Net152 network without mobility learning, the recognition accuracy is improved by 20.9%.

By improving the convolution center size of the first group of the network and adjusting the order of "Convolution Group (Conv) + Batch Normalization Group (BN) + Unconventionality Starting Function Group (Relu)" in the residual module to improve the network, the accuracy of the style recognition is greatly improved. The performance of the network is enhanced. Fig. 7(c) shows the loss function and recognition accuracy of the Res-Net152 network using mobility learning for training and optimization on a collection of girl’s clothing advertisements array. The curve shows that when this network is trained and optimized on a collection of small-volume arrays, the accuracy is the highest at the 68th iteration, which is 92.9%, and the corresponding loss function is 0.134. Compared with the Res-Net152 network without mobility learning, the loss function is 0.134. degree learning, the recognition accuracy is improved by 27.9%. Migrating the model representation parameters from the Image-Net image array collection to the Res-Net152 network reduces the consumption of computational resources and significantly improves the style recognition accuracy. Fig. 7(d) shows the loss function and recognition accuracy of the improved Res-Net152 network using mobility learning, which is trained and optimized on a collection of girl’s clothing advertisement arrays. The graph shows that this network has the highest accuracy of 94.2% at the 66th iteration with a loss function of 0.093 when it is trained on a collection of small batch arrays, compared with the recognition accuracy of the Res-Net152 network without mobility learning, the improved Res-Net152 network without mobility learning, and the Res-Net152 network with mobility learning. Compared with the recognition accuracy of the network, the improved Res-Net152 network using mobility learning is the best in recognizing styles in the collection of girl’s clothing advertisement arrays. Its recognition accuracy is improved by 29.2%, 8.3%, and 1.3%, respectively. This demonstrates the efficiency of the proposed method in this paper in the task of style recognition of girls’ clothing advertisements.

Fig. 7. VAL and Loss of network structure (a) Res-Net152 without using mobility learning; (b) Improved Res-Net152 without using mobility learning; (c) Res-Net152 using mobility learning; (d) Improved Res-Net152 using mobility learning.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig7.png

Through the method of this paper, we can train deep learning models on the color, shape, texture and other features of advertisement images, so as to achieve the classification, recognition and design of advertisement images, and help advertisement platforms to achieve the accurate placement of advertisements. To further test the reliability of the results, we use the improved Res-Net152 network with migration degree learning to test the application of the degree of accurate placement of advertisements. Among them, the efficiency of the advertisement images designed by this system is compared, counted and organized by comparing with the existing related research studies. The results of the judging criteria (acc, P, R and F1) for the above 9 test data (A-I) are shown in Fig. 8.

Fig. 8. Practical results of the test data.

../../Resources/ieie/IEIESPC.2026.15.4.478/fig8.png

It can be seen that the accuracy of placing ads is in the range of 74.0% to 86.8% and the precision is in the interval range of 77% to 91.9%. In recall the interval range is from 79% to 91.9%. the F1 value is stable from 83.7% to 90.8%. It can be seen from the above results that the results of this test have a good advantage in the identification, application and targeted placement of advertisements, and can well capture the accuracy and precision of advertisement placement. Of course, since the process of using this paper’s method to carry out the process from identification, design to targeted placement of ads contains three steps, there are still some errors. However, to summarize, the method in this paper has strong practical application value in the field of advertising.

6. Discussion

A method for recognizing the image style of girls’ clothing advertisements based on improved Res-Net152 network and transfer learning is proposed. The specific conclusions are as follows.

The experimental results show that on the girls’ clothing advertisement dataset, the method achieves the highest recognition accuracy of 94.2% at the 66th iteration with a loss of 0.093, which improves the recognition effect compared with other methods. The method is informative for the study of digitization of girls’ clothing advertisements.

Despite the good results of the method, there are still shortcomings, such as fewer advertisement style types and limited sample size of the dataset. Future work will increase the style types and sample size to further optimize the model.

In practical application, the method performs well in the accuracy, precision and recall of advertisement placement, and the F1 value is stable from 83.7% to 90.8%, which indicates that the method has practical application value in the field of advertising.

The proposed migration learning recognition method based on the improved Res-Net 152 network in this study has achieved significant results on girl’s clothing advertisement data, and its optimized multi-scale feature learning capability and the strategy of accelerating the convergence of the pre-trained model using image nets show that the method has a good generalization potential, which is expected to be applicable to other related fields. The specific code distinction needs to adjust the network structure and parameters according to the actual task.

Funding

2020 Liaoning Provincial Department of Education Science Research Fund Project (Humanities and Social Sciences). No. JFW202015402.

References

1 
H. Zhao , Y. Zhang , G. Liu , Research on fine-grained image classification algorithm based on RPN and B-network architecture, Computer Applications and Software, Vol. 36, No. 3, pp. 210-213, 2019Google Search
2 
X. Zhong , Research and implementation of autonomous classification of clothing styles based on cognitive features, M.S. thesis, Donghua University, Shanghai, China, 2012Google Search
3 
X. Chao , M. J. Huiskes , T. Gritti , C. Ciuhu , A framework for robust feature selection for real-time fashion style recommendation, Proceedings of the 1st International Workshop on Interactive Multimedia for Consumer Electronics, pp. 35-41, 2009Google Search
4 
L. F. Chuang , J. W. Lin , Recognizing and classifying the degree of structural features of clothing advertisement styles based on improved Canny algorithm, Laboratory Research and Exploration, Vol. 39, No. 5, pp. 264-268, 2020Google Search
5 
Y. Gao , B. Wang , Z. Guo , Research on the classification algorithm for apparel image recognition with improved HSR-FCN, Computer Engineering and Applications, Vol. 55, No. 16, pp. 144-149, 2019Google Search
6 
K. He , X. Zhang , S. Ren , J. Sun , Deep residual learning for image recognition, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016DOI
7 
M. Koziarski , B. Cyganek , Image recognition with deep neural networks in presence of noise–Dealing with and taking advantage of distortions, Integrated Computer-Aided Engineering, Vol. 24, No. 4, pp. 337-349, 2017DOI
8 
S. Shubathra , P. C. D. Kalaivaani , S. Santhoshkumar , Clothing image recognition based on multiple features using deep neural networks, 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC), pp. 166-172, 2020DOI
9 
J. Wang , Y. Fan , Z. Li , Texture image recognition based on deep convolutional neural networks and transfer learning, Journal of Computer-Aided Design & Computer Graphics, Vol. 34, No. 5, pp. 701-710, 2022DOI
10 
M. Elleuch , A. Mezghani , M. Khemakhem , M. Kherallah , Clothing classification using deep CNN architecture based on transfer learning, Hybrid Intelligent Systems: Proceedings of the 19th International Conference on Hybrid Intelligent Systems (HIS 2019), Advances in Intelligent Systems and Computing, Vol. 1179, pp. 240-248, 2021DOI
11 
T.-Y. Lin , A. RoyChowdhury , S. Maji , Bilinear CNN models for fine-grained visual recognition, 2015 IEEE International Conference on Computer Vision (ICCV), pp. 1449-1457, 2015DOI
12 
D. Han , Q. Liu , W. Fan , A new image classification method using CNN transfer learning and web data augmentation, Expert Systems with Applications, Vol. 95, pp. 43-56, 2018DOI
13 
S. J. Pan , Q. Yang , A survey on transfer learning, IEEE Transactions on Knowledge and Data Engineering, Vol. 22, No. 10, pp. 1345-1359, 2010DOI
14 
Y. Bengio , Deep learning of representations for unsupervised and transfer learning, Proceedings of the ICML Workshop on Unsupervised and Transfer Learning, JMLR Workshop and Conference Proceedings, Vol. 27, pp. 17-36, 2012Google Search
15 
C. Tan , F. Sun , T. Kong , W. Zhang , C. Yang , C. Liu , A survey on deep transfer learning, Artificial Neural Networks and Machine Learning–ICANN 2018, Lecture Notes in Computer Science, Vol. 11141, pp. 270-279, 2018Google Search
16 
S. Bhattacharyya , A brief survey of color image preprocessing and segmentation techniques, Journal of Pattern Recognition Research, Vol. 6, No. 1, pp. 120-129, 2011DOI
17 
P. Mishra , A. Biancolillo , J. M. Roger , F. Marini , D. N. Rutledge , New data preprocessing trends based on an ensemble of multiple preprocessing techniques, TrAC Trends in Analytical Chemistry, Vol. 132, Art. no. 116045, 2020DOI
18 
C. Shorten , T. M. Khoshgoftaar , A survey on image data augmentation for deep learning, Journal of Big Data, Vol. 6, Art. no. 60, 2019DOI
19 
T. Schlett , C. Rathgeb , C. Busch , Deep learning-based single image face depth data enhancement, Computer Vision and Image Understanding, Vol. 210, Art. no. 103247, 2021DOI
20 
A. Karpathy , G. Toderici , S. Shetty , T. Leung , R. Sukthankar , L. Fei-Fei , Large-scale video classification with convolutional neural networks, 2014 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1725-1732, 2014DOI
21 
J. Yang , K. Yu , Y. Gong , T. Huang , Linear spatial pyramid matching using sparse coding for image classification, 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1794-1801, 2009DOI
22 
J. Li , J.-H. Cheng , J.-Y. Shi , F. Huang , Brief introduction of back propagation (BP) neural network algorithm and its improvement, Advances in Computer Science and Information Engineering, Advances in Intelligent and Soft Computing, Vol. 169, pp. 553-558, 2012DOI
23 
G. Gkioxari , R. Girshick , J. Malik , Contextual action recognition with R*CNN, 2015 IEEE International Conference on Computer Vision (ICCV), pp. 1080-1088, 2015DOI
24 
R. Girshick , Fast R-CNN, 2015 IEEE International Conference on Computer Vision (ICCV), pp. 1440-1448, 2015DOI
25 
S. Ren , K. He , R. Girshick , J. Sun , Faster R-CNN: Towards real-time object detection with region proposal networks, Advances in Neural Information Processing Systems, Vol. 28, pp. 91-99, 2015DOI
26 
J. Redmon , S. Divvala , R. Girshick , A. Farhadi , You only look once: Unified, real-time object detection, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 779-788, 2016DOI
27 
J. Redmon , A. Farhadi , YOLO9000: Better, faster, stronger, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7263-7271, 2017DOI
28 
A. Wong , M. J. Shafiee , F. Li , B. Chwyl , Tiny SSD: A tiny single-shot detection deep convolutional neural network for real-time embedded object detection, arXiv preprint arXiv:1802.06488, 2018DOI
29 
Y. Ma , Status quo and development trend of China's online advertising market, Today's Media, No. 2, pp. 84-86, 2009Google Search
ChunYing Song
../../Resources/ieie/IEIESPC.2026.15.4.478/au1.png

ChunYing Song received a bachelor’s degree in economics from Liaoning University in 2002 and her master’s degree in communication from the school of culture and media, Liaoning University in 2005. She is currently working as an associate professor of advertising at the school of culture, media and art design, Liaoning University of technology.The main research fields and directions include: brand economy, AI advertising and media vision, and advertising development history.

Jian Liu
../../Resources/ieie/IEIESPC.2026.15.4.478/au2.png

Jian Liu received the master degree in in engineering from Liaoning Technical University in 2019. His research areas and directions include computer vision, deep learning and graphic image processing.