Open Access
ARTICLE
Enhancing Personalized Fashion Recommendation by Integrating Large Language Models with Attribute Features
1 Department of Computer Science and Information Engineering, Chang Gung University, Taoyuan, Taiwan
2 Department of Artificial Intelligence, Chang Gung University, Taoyuan, Taiwan
3 Center for Artificial Intelligence in Medicine, Chang Gung Memorial Hospital at Linkou, Taoyuan, Taiwan
* Corresponding Author: Hsien-Tsung Chang. Email:
(This article belongs to the Special Issue: Advances in Natural Language Processing and Large-scale AI Models)
Computer Modeling in Engineering & Sciences 2026, 148(2), 31 https://doi.org/10.32604/cmes.2026.086762
Received 05 June 2026; Accepted 12 August 2026; Issue published 28 August 2026
Abstract
Personalized fashion recommendation requires models that can capture visual compatibility, textual semantics, structured attributes, and user-specific preferences. However, existing multimodal approaches often rely on static word embeddings and shallow text encoders, limiting their ability to represent nuanced fashion descriptions. This study proposes a multimodal recommendation framework enhanced by large language models (LLMs) that integrates visual features, contextual textual representations, and structured attribute features for personalized outfit matching. A Japanese pretrained BERT encoder is used to replace the conventional Word2Vec and convolutional neural network (CNN)-based text pipeline, while GPT-4o is employed to extract fine-grained fashion attributes from product metadata. In addition, Llama-3.3-70B-Instruct is used to estimate semantic similarity among attribute values, enabling attribute-aware compatibility modeling beyond exact matching. Experiments on the IQON3000 dataset, containing 216,791 top-bottom outfit combinations, show that the proposed model achieves an area under the receiver operating characteristic curve (AUC) of 0.8477, outperforming the original Personalized Outfit Recommendation Scheme with Attribute-wise Interpretability based on Bayesian Personalized Ranking (PAI-BPR) baseline. Ranking-based evaluation further demonstrates that the proposed model consistently outperforms PAI-BPR across different candidate-set sizes and places compatible items closer to the top of the recommendation list. These results demonstrate that integrating large language models with multimodal and structured attribute features can effectively improve the accuracy and personalization of fashion recommendation systems.Keywords
In recent years, one of the most revolutionary breakthroughs in artificial intelligence has been the emergence of Large Language Models (LLMs) [1]. From early rule-based and statistical approaches in natural language processing (NLP) to the subsequent development of neural architectures such as Recurrent Neural Networks (RNNs) [2] and Convolutional Neural Networks (CNNs) [3], and now to large-scale pretraining based on the Transformer architecture [4], language understanding technologies have experienced a qualitative leap. Through large-scale pretraining on massive text corpora, LLMs are able to capture not only linguistic structures but also general world knowledge and reasoning patterns. Compared with task-specific models, LLMs exhibit superior generalization and transfer capabilities, allowing them to be flexible to a wide variety of downstream applications. This combination of broad linguistic competence and contextual reasoning has positioned LLMs as key enablers in diverse domains.
Within the domain of e-commerce and recommender systems, LLMs have demonstrated exceptional text understanding abilities that enable the deep analysis of product descriptions and user reviews. In particular, the fashion industry provides an ideal testbed for exploring LLM capabilities, as fashion-related texts often carry multi-layered meanings involving aesthetics, emotions, and cultural context. LLMs can accurately interpret complex expressions such as “minimalist chic,” “vintage style,” or “elegant temperament,” and capture subtle stylistic cues and affective tones that go beyond surface lexical similarity. This capacity for nuanced comprehension makes LLMs promising tools for fine-grained product analysis and personalized fashion recommendation, where understanding both content and style is critical.
Despite these advantages, most existing multimodal recommendation systems continue to rely on traditional text feature extraction techniques—such as Word2Vec [5] embeddings combined with CNN encoders—that are limited in their ability to capture deep semantic structures. These conventional approaches often perform adequately for general text classification but struggle when applied to domains with highly specialized or stylistically rich language, such as fashion. Fashion product descriptions typically include information across multiple semantic layers: material (e.g., “soft silk,” “stretch cotton”), style (e.g., “French elegance,” “street fashion”), usage scenario (e.g., “business formal,” “casual vacation”), and even aesthetic tone or emotional impression (e.g., “intellectual charm,” “youthful vitality”). When a system must recommend a lower-body garment to complement an upper-body item such as a “gentle style knitted top,” it must be able to infer the underlying esthetic concept conveyed by “gentle style” and identify stylistically compatible options, such as a flowing chiffon skirt or high-waisted wide-leg pants. These expressions frequently incorporate cultural and subjective interpretations—terms such as “gentle style” or “salt-based look” represent not just literal meanings but cultural semantics specific to fashion that conventional embeddings do not capture.
With the rapid advancement of LLM technologies—exemplified by models such as BERT [6], GPT [7] and LLaMA [8]—it has become feasible to represent text with rich context-dependent semantics that encompass stylistic, emotional and cultural dimensions. This opens up new opportunities for leveraging LLMs in fashion recommendation tasks, enabling systems to move from shallow word-level similarity toward deeper semantic and conceptual understanding. The central research question of this work is therefore: How can the semantic understanding power of LLMs be harnessed to more accurately analyze fashion descriptions and generate contextually appropriate outfit recommendations? Addressing this question requires exploring how different LLM architectures can be effectively integrated into a multimodal recommendation framework and assessing whether these integrations yield measurable improvements in recommendation quality.
To achieve this goal, this study proposes a multimodal fashion recommendation framework that incorporates large language models as advanced text feature extractors. Given a user ID and an upper-body garment (e.g., a T-shirt, shirt, or jacket), the system automatically recommends the most compatible lower-body garment (e.g., pants or skirts) by jointly considering visual compatibility (color harmony, material coherence, and style alignment) and individual user preferences inferred from historical interactions. Fig. 1 illustrates the overall architecture of the system. The framework is designed to integrate LLM-derived textual representations with visual and attribute-level features, forming a tri-modal system that exploits the complementary strengths of each modality. Specifically, GPT-4o is employed to extract fine-grained item attributes—including color, design, pattern, sleeve length, garment length, silhouette, sleeve type, and neckline—from product metadata. LLaMA is then used to estimate semantic similarity between attribute values within the same category. These structured attributes and similarity scores are incorporated into a dedicated attribute module that complements the textual and visual representations.

Figure 1: Overall architecture of the proposed LLM-integrated multimodal fashion recommendation system.
Building upon this integration, the objective of this study is to substantially enhance the semantic interpretability and personalization capability of fashion recommendation systems. Using the deep semantic reasoning power of LLMs, the proposed model advances from surface-level lexical matching to concept-level understanding, allowing it to better capture the nuanced semantics of fashion language. The effectiveness of the proposed method is validated in the IQON3000 data set, demonstrating its ability to outperform the baseline multimodal models in outfit recommendation performance.
The contributions of this work are twofold:
1. Deep semantic text feature extraction: This study introduces BERT as a pretrained language encoder to replace traditional Word2Vec and CNN-based text processing methods, upgrading from static 300-dimensional word embeddings to dynamic 768-dimensional contextual representations. This transition significantly improves the system’s comprehension of semantic and stylistic nuances in fashion descriptions.
2. Intelligent attribute analysis and modular design: By employing GPT-4o and LLaMA for intelligent attribute analysis, the system extracts fine-grained attributes—including color, design, pattern, sleeve length, garment length, silhouette, sleeve type, and neckline—and computes semantic similarity scores between attribute values. A dedicated attribute module is then incorporated as a third modality, complementing the textual and visual representations to improve recommendation accuracy.
Through this integration of large language models and multimodal learning, this study aims to advance the precision, interpretability, and personalization of fashion recommendation systems, contributing to the broader exploration of how LLMs can serve as semantic engines in complex, style-driven domains.
This section provides a concise overview of research on multimodal recommendation with a focus on fashion. We first trace the evolution of fusion strategies—from early feature concatenation to mid-level interaction and interpretable, attribute-aware models. We then review text representation advances from Term Frequency–Inverse Document Frequency (TF–IDF) and Word2Vec to contextual encoders (ELMo/BERT) and, most recently, large language models (LLMs) used for recommendation. Finally, we highlight practical issues such as cross-modal alignment, efficiency, and reliability (e.g., latency and hallucinations) that remain open. These gaps motivate our LLM-enhanced, attribute-aware multimodal framework for outfit compatibility and personalized fashion recommendation.
2.1 Architectures and Evolution of Multimodal Systems
A central challenge in multimodal recommendation is how to effectively fuse heterogeneous signals (e.g., text, images, and structured metadata) so as to capture complementary semantics and improve expressiveness and predictive performance. Early fusion strategies concatenated features at the input or representation level. A seminal example is VBPR [9], which incorporated CNN-based visual features of fashion items into a Bayesian personalized ranking framework to better model users’ visual preferences. To move beyond simple concatenation, Tensor Fusion Networks (TFN) [10] explicitly modeled higher-order inter-modal interactions via outer products, although at a considerable computational cost; Low-rank Multimodal Fusion (LMF) [11] subsequently reduced this cost through low-rank factorization.
Mid-level (or “bottleneck”) fusion mechanisms further improved the trade-off between accuracy and efficiency by enabling controlled cross-modal exchange while preserving modality-specific processing. Multimodal Bottleneck Transformers (MBT) [12] restrict information exchange to a small set of bottleneck tokens, which has proven effective when jointly modeling visual details from images and stylistic cues from text. Building on these trends, PAI-BPR [13] advances personalized and interpretable fashion recommendation by introducing an Attribute Classification Network (ACN) that decomposes items into fine-grained, human-interpretable attributes (texture, style, material, shape, and part-level details). The model uses a 2,048-dimensional attribute vector (plus an 8-dimensional color quantization via
Recent collaborative filtering studies have also explored graph-based architectures to better capture user–item interaction structures. Alshareet and Ben Hamza [14] proposed an adaptive spectral graph wavelet framework for collaborative filtering, where users, items, and their interactions are represented as a bipartite graph. Their method applies an adaptive transfer function and spectral graph wavelets to learn low-dimensional user and item embeddings, allowing the model to capture both local and global graph structures in implicit-feedback recommendation.
This line of work is relevant because it demonstrates the effectiveness of graph-spectral modeling for collaborative filtering. However, its focus differs from the objective of the present study. Spectral graph wavelet methods mainly improve the modeling of user–item interaction topology, whereas our proposed LLM-PAI-BPR framework focuses on personalized top–bottom fashion compatibility prediction by enriching item representations with visual features, contextual textual representations, and LLM-assisted structured attribute features. Therefore, graph-spectral collaborative filtering and multimodal attribute-enhanced fashion recommendation address different but complementary aspects of recommender-system design.
2.2 Characteristics and Challenges of Fashion Recommendation
Fashion recommendation differs from general product recommendation in its heavy reliance on image and text signals that are difficult to discretize with a small set of tags. Beyond item names, critical details include silhouette, cut, color, and fabric. Type-aware embeddings [15] learn distinct subspaces for item categories (e.g., tops, bottoms, shoes) to model outfit compatibility across types. Scenario-dependent preferences (e.g., work, travel, or dating) further require recommendation strategies to adapt to context; few-shot preference modeling [16] helps for cold-start users and newly listed items. Rapidly evolving terminology and trends pose an additional challenge: emerging expressions (e.g., culturally grounded style descriptors) are hard to capture with traditional word embeddings or shallow CNN-based encoders, which limits semantic coverage and reduces recommendation fidelity.
2.3 Evolution of Text Feature Learning
Classical text representations in recommendation relied on frequency-based models such as TF–IDF [17], which ignore context and inter-word semantics. Word2Vec [5] advanced distributional semantics via CBOW/Skip-gram but yields static embeddings that cannot disambiguate word senses across contexts. Deep models then leveraged pretrained embeddings with convolutional neural network (CNN) and bidirectional long short-term memory (BiLSTM) encoders; for instance, Kim [3] demonstrated that CNNs capture salient local
2.4 Large Language Models for Recommendation
LLMs such as GPT and BERT have recently attracted attention in recommendation for their superior semantic understanding of user reviews and product descriptions. Empirical studies report gains from integrating unstructured signals (e.g., social media and review text) into recommenders [19]. To address the latency and deployment costs of large models, knowledge distillation to smaller student models (e.g., SLMRec) has been explored to retain semantic competence while improving efficiency [20]. In multimodal settings, LLMs can act as cross-modal “translators” to alleviate semantic misalignment between vision and language, which is particularly valuable in fashion tasks requiring both visual reasoning and stylistic understanding [21]. Nonetheless, practical issues remain: model size can induce serving delays and, in some cases, hallucinated recommendations; recent work highlights the need for efficiency, reliability, and alignment in LLM-driven recommenders [22].
2.5 Benchmark Evidence of LLM Understanding
Because this study leverages LLMs to process product descriptions and user reviews, it is important to situate their language understanding against standardized benchmarks. BERT, for example, reports strong scores on the General Language Understanding Evaluation (GLUE) benchmark and achieves high F1 on SQuAD v1.1 [6], while GPT-4 and successors attain competitive or near-expert performance on comprehensive benchmarks such as MMLU [23,24]. Although benchmark outcomes do not directly translate to recommendation quality, they indicate that modern LLMs possess robust contextual reasoning and semantic competence, motivating their adoption as textual backbones within multimodal fashion recommendation frameworks.
This section describes in detail the methodology and techniques employed in this study to enhance the precision of personalized fashion recommendations. We propose a recommendation framework that deeply integrates Large Language Models (LLMs) with a tri-modal feature engineering approach. The section first introduces the dataset used in this study and explains how it is divided into training, validation, and test subsets. We then describe how visual, textual, and attribute-based features are extracted and processed, with particular emphasis on the use of LLMs—such as GPT-4o, LLaMA, and BERT—for understanding textual semantics, extracting product attributes, and computing attribute similarities. Finally, we explain the core architecture of our recommendation model, which jointly considers general outfit compatibility and individual user preferences. Fig. 2 presents the overall architecture of the proposed recommendation system, detailing the complete processing pipeline from raw data input to final recommendation output.

Figure 2: Overall system architecture of the proposed recommendation framework. The pipeline integrates visual feature extraction, BERT-based textual encoding, GPT-4o-based attribute extraction, LLaMA-based attribute similarity modeling, feature fusion, and personal preference modeling.
3.1 Multimodal Feature Engineering
We adopt a tri-modal fusion architecture comprising visual, textual, and attribute features. For the visual modality, following PAI-BPR [13], we employ a modified AlexNet trained for multi-task attribute classification to produce a 2,048-dimensional visual semantic vector.
We evaluate our approach on the IQON3000 dataset [25], which is curated from the Japanese fashion social platform IQON. After collection and preprocessing by the original authors— including duplicate removal and quality control—the dataset contains 216,791 top–bottom outfit pairs created by 3568 users. IQON3000 focuses specifically on the outfit recommendation task between tops and bottoms, excluding accessories, shoes, and other categories. For each outfit, the data include item images and a structured JSON metadata file that records multi-dimensional attributes, such as itemName, categorys, hierarchical breadcrumb, colors, product options, and stylistic expressions. These rich textual signals provide a strong semantic basis for our BERT-based feature extraction and downstream LLM-driven analyses.
We adopt the standard train–validation–test split used in prior work. The detailed configuration is shown in Table 1.

3.1.2 Selection and Application of Large Language Models
This study leverages three large language models at different stages of the pipeline, each serving a distinct role and exploiting complementary strengths.
We employ GPT-4o [26] for attribute analysis on the IQON3000 dataset, primarily because of its strong multimodal understanding and extensive domain knowledge in fashion. By jointly processing images and textual metadata, GPT-4o identifies fine-grained item attributes, including color, design, pattern, sleeve length, garment length, silhouette, sleeve type, and neckline, and organizes them into a structured attribute taxonomy.
To construct an attribute–similarity lookup table, which requires a large number of application programming interface (API) calls (involving similarity judgments over thousands of attribute-value pairs), we adopt Meta’s Llama-3.3-70B-Instruct [27] as a cost-effective solution. This model maintains high-quality semantic understanding while being more amenable to large-scale batch processing. Our experiments indicate that Llama-3.3-70B delivers competitive performance for fashion-domain semantic similarity estimation, supporting semantic smoothing of discrete attribute embeddings in the downstream recommendation model.
In our implementation, the similarity scoring prompt is fixed, and the model outputs a numerical score in the range
For Japanese text encoding, we replace the original CNN+Word2Vec pipeline with tohoku-nlp/bert-base-japanese [28]. Pretrained specifically for Japanese, this BERT variant better handles the product descriptions in IQON3000 and, through bidirectional attention, captures richer contextual semantics. This BERT encoder is used as a frozen feature extractor rather than being fine-tuned during recommendation model training.
Taken together, this multi-model strategy exploits the complementary strengths of different LLMs—supporting robust attribute extraction, semantic similarity evaluation, and high-quality text understanding across the key components of our system.
3.1.3 Textual Feature Enhancement
We replace the conventional CNN+Word2Vec pipeline with a Transformer-based BERT encoder for textual feature processing. Unlike Text-CNN, which relies on multiple convolutional kernels and pooling operations, BERT’s bidirectional contextual modeling captures sentence-level semantics more comprehensively. This substitution not only improves the quality of semantic representations but also simplifies the overall architecture, avoiding additional convolution/pooling layers and easing integration into the recommendation model.
All product-related text in the IQON3000 dataset is written in Japanese—covering fields such as item names, brand information, and category descriptions. In the original JSON files, Japanese strings are stored as Unicode escape sequences (e.g., ∖u30b8∖u30e3∖u30b1∖u30c3∖u30c8 represents “ジャケット”), which decode to standard Japanese text during parsing. To accommodate the linguistic characteristics of Japanese (e.g., mixed usage of kana and kanji, relatively flexible word order), we adopt tohoku-nlp/bert-base-japanese as our text encoder. Pretrained on Japanese corpora, this model provides more accurate handling of domain-specific fashion vocabulary than multilingual BERT, enabling finer-grained understanding of nuances in product descriptions.
For feature construction, we concatenate multiple textual fields—itemName, categorys, colors, options, and expressions—after trimming whitespace and removing duplicated phrases, forming a unified semantic description per item. The combined text is then encoded by BERT, and the output of the [CLS] token is used as the item’s textual semantic representation, yielding a 768-dimensional dense vector. The overall process is illustrated in Fig. 3.

Figure 3: Text feature processing pipeline. Product metadata fields are consolidated into a unified textual description and encoded by a Japanese BERT model to generate a 768-dimensional semantic representation.
3.1.4 Attribute Feature Design
While BERT excels at contextual semantic understanding, the JSON metadata of each product contains numerous discrete attribute tags—such as color codes, size specifications, and style categories. When these structured attributes are verbalized into free-form text, their discrete nature and clear categorical boundaries may be diluted in BERT’s continuous semantic space. To preserve and exploit this information, we design a dedicated attribute module that complements BERT-based textual features: (i) when product descriptions are short or missing, structured attributes provide a reliable fallback representation; and (ii) the explicit categorical structure of attributes improves interpretability, enabling the system to articulate recommendation rationales (e.g., “same color family,” “style consistency”).
Conventional attribute matching typically checks only for exact equality (e.g., “red” equals “red”) and cannot capture fine-grained semantic relations such as the proximity between “pink” and “light pink.” We therefore introduce an LLM-based attribute similarity computation to model nuanced relations among attribute values. This allows the system to determine which attributes are compatible even when they are not identical, effectively bridging structured attribute signals with semantic understanding and thereby improving both accuracy and flexibility in recommendation. The overall attribute feature processing pipeline is illustrated in Fig. 4. It consists of GPT-4o-based attribute extraction, attribute vector encoding, and LLaMA-based attribute similarity table construction.

Figure 4: Attribute feature processing pipeline. Product metadata are first converted into structured attribute JSON outputs using GPT-4o, then encoded into attribute vectors and further used to construct an attribute similarity mapping table with LLaMA.
Attribute Extraction and Categorization
To convert unstructured product descriptions into computable structured representations, we employ OpenAI’s GPT-4o Large Language Model (LLM) to perform semantic parsing and attribute extraction over the JSON metadata. We design a structured prompt template—comprising a task description, definitions of eight attribute categories, output formatting rules, and illustrative examples (see full prompt in Fig. 5)—to guide the model toward consistent, schema-aligned outputs.

Figure 5: English-translated prompt template used for fashion item classification and high-level attribute extraction. In the actual implementation, Chinese task instructions were used together with the original Japanese IQON3000 product metadata. The model was instructed to return only a valid JavaScript Object Notation (JSON) object for downstream processing.
Leveraging GPT-4o’s deep language understanding, we extract key fashion attributes and normalize them into (category, value) pairs under eight high-level categories: Color, Design, Pattern, Sleeve Length, Length, Silhouette, Sleeve, and Neck. The resulting vocabulary sizes for each category are as follows: Color (257), Design (214), Pattern (254), Sleeve Length (315), Length (69), Silhouette (58), Sleeve (216), and Neck (34), totaling 1417 fine-grained attribute values.
The attribute vocabulary and the LLaMA-based attribute similarity table are constructed from the full processed IQON3000 item catalog, including items appearing in the training, validation, and test splits. Therefore, the evaluation follows a transductive item-catalog setting rather than an inductive cold-start setting. Test-set interaction labels and outfit-pair compatibility labels are not used for model training, early stopping, or modality-weight selection; they are used only for final evaluation.
To improve extraction consistency, GPT-4o is used with a fixed prompt, JSON-format output constraints, and deterministic decoding with temperature set to 0. The extracted results are cached so that each item is processed consistently across experiments, and invalid or missing attributes are mapped to the reserved missing-value index in the attribute module.
These controls improve consistency but do not replace an independent validation of extraction quality. Since IQON3000 does not provide ground-truth labels for the eight normalized attribute categories, this study does not report per-category precision, recall, or F1-score for GPT-4o-based attribute extraction. Manual or semi-automatic validation of schema violations, unsupported attributes, missing rates, and consistency with the original Japanese metadata remains future work.
The distribution of the extracted attribute values and an example of the structured JSON output are shown in Fig. 6.

Figure 6: Attribute distribution statistics and an example of product attribute extraction. The left panel reports the value counts of the eight fashion attribute categories, and the right panel presents an example structured JSON output generated by the attribute extraction process.
Attribute Encoding and Vector Representation
To obtain a numeric representation of product attributes, we assign a unique positive integer ID to each attribute value, yielding a compact encoding scheme. For example, the encoding map for the Color category is shown in Table 2.

Based on this scheme, each item is represented by an 8-dimensional attribute vector
As an illustrative example, Table 3 shows that item ID 37396155 is represented by the following 8-dimensional attribute vector:
where the elements correspond to Color, Design, Pattern, Sleeve Length, Length, Silhouette, Sleeve, and Neck, respectively; a value of

Through this process, IQON3000 items are transformed from unstructured fashion descriptions into standardized numerical representations, providing a consistent data foundation for model training, fashion recommendation algorithms, and style similarity computation.
Naively computing similarity via inner products is fundamentally flawed in our setting. The attribute IDs are assigned by enumerating the GPT-4o–derived attribute lists within each category on IQON3000; hence, the numeric IDs themselves carry no semantic meaning. For example, within the Color category, “Pink” might be encoded as 1 and “Baby Pink” as 11. Although these two colors are close in colorimetry, the numeric gap between their IDs does not reflect semantic proximity, causing inner-product–based similarity to deviate markedly from true attribute similarity.
To address this issue, we construct a per-category similarity lookup table over all pairs of attribute values using Meta’s Llama-3.3-70B-Instruct. We design a dedicated prompt (see Fig. 7) that instructs the model to assess pairwise similarity grounded in fashion knowledge and to output a score in

Figure 7: English-translated prompt template used for attribute similarity scoring. In the actual implementation, Chinese task instructions were used to ask Llama-3.3-70B-Instruct to evaluate whether two normalized attribute values are semantically similar or interchangeable within the same fashion attribute category. The model was instructed to return only a single numerical score in the range of

Figure 8: Example of an attribute similarity mapping table for the Color category generated by Llama-3.3-70B-Instruct. Each key represents a pair of attribute values, and each score indicates their semantic similarity for outfit coordination.
For GPT-4o-based attribute extraction, the model identifier is gpt-4o. The temperature is set to 0.0 because the task requires stable JSON-formatted outputs rather than diverse generation. The prompt explicitly asks the model to return only a JSON object, and the extracted results are cached for reuse. The top_p value and maximum output length are not explicitly specified in the original API call; therefore, the API default settings are used.
For the LLaMA-based attribute similarity table, the model identifier is meta-llama/Llama-3.3-70B-Instruct. The generation settings are temperature=0.3, top_p=0.8, and max_tokens=10. This low-temperature nucleus-sampling setting is used to obtain relatively stable numerical similarity scores while keeping the output short. The prompt asks the model to return only a numerical score in the range
During preprocessing, only one score is generated for each unordered pair of attribute values. Therefore, reverse pairs are not queried separately, and the same stored score is used for both directions during lookup. The self-similarity of an attribute value is defined as 1.
3.2 Model Architecture and Design
Before presenting the mathematical formulation of the proposed model, Table 4 summarizes the major symbols used throughout this section for ease of reference.

We adopt an LLM-enhanced multimodal fashion recommendation system that models both general compatibility and personal preference in a dual-objective manner [13]. The core formulation is:
where:
•
•
•
•
3.2.2 Tri-Modal Feature Fusion Architecture
We propose a tri-modal fusion mechanism:
where
For the visual modality, we reuse the pre-extracted features from PAI-BPR [13]. These features are obtained by a modified AlexNet trained for multi-task attribute classification and combined with
where
The textual and attribute modalities are also transformed into 512-dimensional representations, as described in the following subsections. Therefore, for each top–bottom pair
The similarity scores in the visual, textual, and attribute modalities are computed using these aligned 512-dimensional representations, ensuring dimensional consistency before the weighted fusion step.
3.2.3 Textual Feature Enhancement Module
For the textual modality, encoding product descriptions with tohoku-nlp/bert-base-japanese yields a 768-dimensional semantic vector. Since the multimodal fusion module operates in a shared 512-dimensional latent space, the BERT-based representation must be transformed before fusion. To align dimensions while preserving important semantic information, we adopt a bottlenecked TextEncoder module:
The first linear transformation compresses the 768-dimensional BERT representation into a 384-dimensional bottleneck representation, followed by a GELU (Gaussian Error Linear Unit) activation and Dropout regularization. The second transformation maps the compressed representation to 512 dimensions, followed by Layer Normalization and Dropout. This bottleneck design encourages the model to select and retain salient semantic cues from the BERT representation while producing a textual feature vector compatible with the visual and attribute modalities.
Formally, given the BERT textual representation
where
3.2.4 Attribute Feature Module
The attribute feature module is tailored to encode explicit semantic properties of products (e.g., color, pattern, design, sleeve type, neckline), complementing visual and textual signals where they may fail to capture fine-grained categorical differences. It further leverages an attribute similarity matrix to increase the recommender’s sensitivity to semantic compatibility.
For each attribute category, we instantiate a trainable embedding layer that maps attribute values to 64-dimensional dense vectors. Index 0 in every embedding table is reserved to denote missingness (e.g., an item without a recorded neckline) and is mapped to the all-zero vector so that the model can explicitly recognize and handle missing attributes—reflecting realistic data conditions.
Because raw attribute labels are discrete and their ordinal IDs do not encode semantics, we introduce a similarity-based smoothing mechanism. Building on an LLM-derived similarity lookup table, we model pairwise semantic proximity among values within each attribute category. After the initial lookup embedding, we compute a similarity-weighted average of embeddings according to the lookup scores, so that semantically close but differently coded values (e.g., “Pink” and “Light Pink”) are pulled nearer in the vector space, improving semantic generalization and compatibility reasoning.
After embedding and smoothing, the eight attribute embeddings are concatenated to form a 512-dimensional preliminary attribute feature vector. This vector is then fed into a three-layer Multi-Layer Perceptron (MLP) with two ReLU nonlinearities and Dropout regularization, finally producing a 512-dimensional attribute semantic representation. This representation serves as the attribute modality in the tri-modal fusion architecture, participating—together with textual and visual semantics—in feature fusion and compatibility estimation for outfit recommendation.
3.2.5 Personal Preference Modeling
This component aims to capture each user’s individualized style tendencies in clothing selection. Concretely, for every user the system learns preferences over different styles or attributes—for example, some users favor light colors, some prefer formal styles, while others particularly like specific patterns. Such preferences may reside in the item’s visual appearance, textual description, or structured attributes (e.g., color, material, pattern). Accordingly, we decompose user preference learning into three modalities: visual, textual, and attribute.
The overall personal preference score
The formulation consists of three parts:
• Bias terms:
• Latent factor interaction:
• Multimodal semantic interaction: For each user, we learn modality-specific preference vectors
This modeling approach uncovers users’ implicit style inclinations while leveraging interpretable semantic cues (e.g., attributes like “striped” or “chiffon”), thereby providing a more comprehensive profile of fashion preferences. The resulting scores are matched against each item’s multimodal features and used for ranking.
3.3 Training Phase and Loss Function
We train the model using the Bayesian Personalized Ranking (BPR) objective [29], whose core goal is to rank items that a user prefers ahead of those the user dislikes. In practice, we follow the proposed PAI-BPR framework [13].
At each training step, the model samples a quadruple
The loss function is defined as:
where
Because our task focuses on relative ranking in outfit recommendation—prioritizing better top–bottom matches rather than merely deciding suitability—we adopt two primary metrics: the Area Under the Receiver Operating Characteristic Curve (AUC) and Mean Reciprocal Rank (MRR). AUC assesses whether preferred matches are ranked ahead of non-preferred ones, while MRR emphasizes the system’s ability to surface the single best match in a retrieval setting. The two are complementary and together capture overall ranking quality and personalization accuracy.
3.4.1 Area under the ROC Curve (AUC)
We select AUC as a principal metric because the core objective in recommender systems is ranking, not binary classification. In our experiments, AUC is computed under the BPR pairwise-comparison scheme, measuring whether a user-preferred outfit
For fairness and comparability, we follow the original PAI-BPR data split protocol. We select the best checkpoint on the validation set and report the final AUC on the held-out test set, mitigating overfitting and ensuring reliable evaluation.
3.4.2 Mean Reciprocal Rank (MRR)
As a complementary metric, we report MRR to quantify ranking performance in a personalized retrieval scenario. Unlike AUC’s pairwise focus, MRR is suited to cases where the system must select the single most appropriate item from a candidate set. Concretely, given a query composed of user
If the correct item is ranked first, the reciprocal rank is 1; if third, it is
3.4.3 Top-
In practical recommender interfaces, users typically only inspect the top few results. Beyond overall MRR, we therefore report MRR@
where
In summary, our methodology strengthens the baseline outfit recommendation pipeline through attribute vector construction, semantic extraction, similarity lookup, and fusion strategies. By combining LLM-based processing of unstructured text with attribute-level compatibility modeling, the system more finely assesses item–item compatibility and user preferences, thereby improving overall recommendation effectiveness. The next section presents experimental results and analysis demonstrating the method’s practical performance.
This section systematically presents and analyzes the experimental results of our multimodal personalized fashion recommendation model. We organize the evaluation into four parts. First, we quantify the contribution of each modality (vision, text, attributes) via single-modality ablations and weighted-fusion studies to identify optimal fusion settings. Second, we compare against multiple baselines to verify the accuracy gains of our approach. Third, we assess ranking performance using Mean Reciprocal Rank (MRR), simulating practical Top-
4.1 Experimental Environment and Settings
This section details the experimental setup and dataset specifications. We first describe the hardware and software environments used for training and evaluation. We then present the datasets, including their sources, scale, and organization.
The hardware setup used in this study is summarized in Table 5.

The software environment is summarized in Table 6, and the key library versions are listed in Table 7.


In multimodal recommendation systems, the contribution of each modality may be uneven. Simply concatenating multimodal features may therefore fail to fully exploit their respective strengths. To further improve performance, we conduct a systematic weight search to identify the optimal combination of Visual (Vis), Text (Text), and Attribute (Attr) feature weights.
All weighting experiments in this section are conducted on the validation set. The training set is used for model parameter learning, the validation set is used for early stopping and modality-weight selection, and the held-out test set is used only for final performance evaluation after the best weight configuration has been fixed.
4.2.1 Unimodal Ablation Analysis
To design a reasonable weighting scheme, we first conduct unimodal ablation experiments in which each modality is fed to the model independently (Table 8). The results show that textual features perform best (AUC = 0.8430), while visual and attribute features are comparable (0.8286 and 0.8271, respectively). This suggests that text serves as the primary discriminative signal, with vision and attributes providing complementary information.

To further examine modality interactions, we also evaluate pairwise fusion settings with equal modality weights. The visual–textual combination achieves the best pairwise result (AUC = 0.8464), outperforming both visual-only and text-only inputs. By contrast, visual–attribute obtains an AUC of 0.8280, and text–attribute reaches 0.8419, slightly below the text-only model.
These results suggest that the attribute modality is useful but weight-sensitive. Since many attributes are extracted from product fields already included in the BERT input, the attribute signal may partially overlap with the textual representation when equally weighted. Therefore, BERT should be regarded as the main contributor to the performance gain, while the GPT-4o/LLaMA-based attribute module serves as a complementary structured feature under appropriate fusion weights.
Given the performance gap across the three modalities, we conduct a systematic search over modality weights to better exploit the potential of the fusion model. We adopt a step size of 0.1 under the constraint
From Table 9, assigning a moderate weight to the textual modality (roughly 0.5–0.7) yields stable and comparatively higher AUC, indicating that text plays a leading role in this task. The best result occurs at Visual=0.3, Text=0.5, Attr=0.2 (AUC = 0.8477). When the text weight becomes too high (e.g., 0.8), performance slightly declines relative to the mid-range settings, and increasing the attribute weight toward 0.4 tends to reduce AUC as well. Overall, these results suggest that a text-centric yet balanced fusion—supplemented by visual cues and a moderate use of attribute signals—provides the most reliable gains.

4.3 Comprehensive Performance Comparison
As shown in Table 10, our method achieves an AUC of 0.8477 on the IQON3000 dataset, outperforming all baselines by a clear margin. Relative to popularity-based approaches (POP-T: 0.6042; POP-U: 0.5951) and the random baseline (RAND: 0.5014), the gains are substantial, underscoring the potential of multimodal deep learning for personalized fashion recommendation. Even against neural baselines such as Bi-LSTM (0.6611) and BPR-DAE (0.6912), which already employ deep architectures, there remains a considerable gap when faced with complex multimodal fashion data.
The importance of multimodal fusion is evident from the comparisons. VTBPR (0.8194) surpasses single-modality variants VBPR (0.8088) and TBPR (0.8102), indicating that jointly leveraging visual and textual information is crucial for capturing user preferences and item characteristics. GP-BPR (0.8321), which simultaneously models item compatibility and user-specific preference, outperforms several earlier baselines that focus on a single objective. However, PAI-BPR achieves a higher AUC of 0.8368, indicating the additional benefit of attribute-aware modeling.
Our improvements highlight the strengths of large language models in understanding and processing natural language text. By introducing BERT-based text encoding together with attribute vectors, the system captures richer, multi-dimensional characteristics of fashion items. BERT handles the semantic complexity of brand information, product descriptions, and category-related text, while the attribute vectors provide structured, discrete representations of item properties. The two components are complementary: the broad linguistic knowledge and strong semantic reasoning of LLMs, coupled with precise attribute-level signals, yield marked performance gains for fashion-text understanding and, consequently, for recommendation accuracy.
4.4 Mean Reciprocal Rank (MRR) Evaluation
To assess ranking performance in the bottom-item pairing task under practical conditions, we adopt Mean Reciprocal Rank (MRR) as a primary metric. MRR is a widely used ranking measure with values in
A higher MRR indicates that the model tends to place the correct bottom item near the top of the list, reflecting better recommendation quality; conversely, a lower MRR suggests weaker ranking capability and a tendency to surface incorrect pairings earlier. This metric is particularly suitable for Top-
We evaluate on a test set comprising 23,095 real top–bottom pairing instances. To compare the model’s ranking ability under varying levels of difficulty, we construct multiple candidate sets per test instance with sizes
• Evaluation size: 23,095 instances.
• Candidates per instance (
• Positive definition: the ground-truth top–bottom pair observed in the original interaction data.
• Negative sampling: bottoms randomly drawn from other users’ inventories, non-duplicated and excluding the positive.
4.4.2 MRR Results and Discussion
As shown in Table 11, our method consistently outperforms the PAI-BPR baseline across all candidate sizes (

With a smaller candidate set (
For both models, MRR naturally decreases as the candidate size increases, since positioning the positive item near the top becomes more difficult with more distractors. Nevertheless, our method remains consistently superior and stable across settings, suggesting good Top-
4.4.3 Top-
To make the evaluation more comprehensive and comparable, we compute MRR@K for all candidate set sizes (
When selecting
• Candidates = 5: MRR@1, MRR@3, MRR@5
• Candidates = 10: MRR@1, MRR@3, MRR@5, MRR@10
• Candidates = 20: MRR@1, MRR@3, MRR@5, MRR@10, MRR@20
• Candidates = 50: MRR@1, MRR@3, MRR@5, MRR@10, MRR@20, MRR@50
This evaluation protocol allows us to observe ranking performance at different Top-
Table 12 reports MRR@K for our method and the PAI-BPR baseline across candidate sizes

As
Performance gains are particularly notable at the top ranks that users most often inspect (e.g., MRR@1 and MRR@3). With
In sum, the proposed method delivers consistent and strong results across candidate sizes and Top-
4.5 Qualitative Results from Real-World Tests
4.5.1 Visualization of Ranking Cases
Fig. 9 illustrates the ranking outcome for a test instance. The query consists of user 429607 with a top item 38031971 and a candidate set of five bottoms. Among the candidates, there is one ground-truth positive bottom 38385533 and four randomly sampled negatives. The model correctly ranks the positive item at the first position, indicating strong discriminative ability at ranking time. The reciprocal rank (RR) for this instance is

Figure 9: Successful ranking example for a top–bottom recommendation query. The ground-truth bottom item is ranked first among five candidate bottoms.
As shown in Fig. 10, this query corresponds to user 2462803 with a top item 11053621 and a candidate set of five bottoms. The objective is to place the positive bottom 11606201 at the top. However, the final ranking places the positive item in the second position. While this is not entirely incorrect and still acceptable in practice (with RR

Figure 10: Failure ranking example for a top–bottom recommendation query. The ground-truth bottom item is ranked second among five candidate bottoms.
This example reflects a common failure type in top–bottom recommendation: near-miss ranking error. In this case, the model does not completely fail to identify a compatible item, since the positive bottom is still ranked within the top two positions. However, the model assigns a slightly higher score to another candidate, indicating that it may have difficulty distinguishing between highly similar or visually plausible alternatives. Such errors are especially important in fashion recommendation because multiple candidate bottoms may appear reasonable from a visual or stylistic perspective, but only one item is treated as the ground-truth positive pair in the evaluation protocol.
This type of error suggests that the proposed model can capture general compatibility signals, but its fine-grained ranking ability remains limited when candidate items share similar visual, textual, or attribute-level characteristics. In particular, the model may overemphasize general style compatibility while failing to sufficiently capture user-specific preference or subtle attribute differences. Therefore, this case indicates that future improvements should focus on more discriminative ranking mechanisms, hard-negative evaluation, and stronger user-preference modeling.
4.5.2 Visualization of Different User Preferences
To further verify whether the system effectively reflects different users’ outfit preferences, we take the same top and generate ranked recommendations for two users with distinct style histories. The results are shown in Fig. 11.

Figure 11: Examples of personalized recommendation results for two users with different historical outfit preferences. Given the same top item, the model generates different bottom-item rankings according to each user’s style history.
The first user’s history favors minimalist styles and denim, with a predominance of slim-fit bottoms. Accordingly, the model places denim candidates (ranks 1–3) at the top of the list, reflecting a preference for basics and practical styles.
By contrast, the second user prefers relaxed silhouettes and floral skirts, with a more casual yet design-forward style history. For the same query top (36235670), the model tends to recommend wide-leg pants and printed skirts (ranks 1–2), yielding a different ranking pattern from the first user’s denim-oriented list.
This case demonstrates that the model adapts the ranking to users’ past outfit preferences, achieving personalized recommendations. It also highlights that the model considers not only inter-item compatibility but also learns and responds to users’ individual styles and tastes.
4.6.1 Effects of System Optimizations
This study successfully leverages the semantic understanding afforded by large language models to improve textual feature processing in the recommender. The AUC increases from 0.8368 to 0.8477, indicating that language models can materially strengthen multimodal recommendation performance. Moreover, the modality-weight analysis identifies the best combination as Text: 0.5, Visual: 0.3, Attr: 0.2, further highlighting the central role of language models in capturing outfit semantics. The MRR results likewise confirm superior ranking accuracy, suggesting tangible gains in user experience and click-through rates.
4.6.2 Computational Cost Discussion
The proposed framework introduces additional computational cost mainly during offline preprocessing. GPT-4o is used for item-level attribute extraction, LLaMA is used to construct the attribute similarity lookup table, and BERT is used to generate textual representations. These large language models are not called during online recommendation inference. In the online stage, the recommender operates on precomputed visual features, precomputed BERT textual features, and encoded attribute representations.
Based on the processed attribute files, the GPT-4o attribute extraction stage covers 128,821 top/bottom items, including 89,095 top items and 39,726 bottom items. Since the extraction process uses caching, the number of GPT-4o calls is bounded by the number of uncached items and is at most one item-level call per processed product.
For the LLaMA-based attribute similarity table, the number of calls is determined by the number of unordered attribute-value pairs within each attribute category:
where
pairwise LLaMA similarity-scoring calls. This table is constructed once offline and then kept fixed during recommendation model training and inference.
The BERT textual features are also precomputed. Since each item is represented by a 768-dimensional float32 vector, storing BERT features for 128,821 items requires approximately 377 MB of storage. Therefore, the main cost-performance trade-off of the proposed framework is that it introduces additional offline preprocessing cost in exchange for improved recommendation performance. Under the adopted IQON3000 evaluation setting, the final model improves the AUC from 0.8368 in PAI-BPR to 0.8477, while avoiding online calls to GPT-4o, LLaMA, or BERT during recommendation inference.
Exact wall-clock preprocessing time, training time, inference latency, and monetary cost were not logged in the original experiments. Therefore, this study reports the preprocessing scale, call-count estimates, and memory footprint rather than reconstructed measured costs. A controlled cost-performance benchmark under the same hardware and software environment remains an important direction for future work.
4.6.3 Limitations and Future Improvements
The current evaluation is based on IQON3000, which is suitable for personalized top–bottom outfit recommendation and allows direct comparison with GP-BPR and PAI-BPR. However, since the dataset mainly contains Japanese product metadata, the results should be interpreted within this dataset setting and should not be viewed as evidence of cross-lingual or cross-cultural generalization. Future work should further evaluate the framework on English-language and culturally diverse fashion datasets with compatible user histories, item metadata, and evaluation protocols.
The experiments are conducted using the standard train–validation–test split of IQON3000 and a fixed experimental setting. Therefore, this study does not report standard deviations, confidence intervals, or statistical significance tests across multiple random seeds. The reported improvements should be interpreted as observed gains under the adopted evaluation protocol, rather than as statistically verified improvements across repeated runs. Future evaluation should include multiple random seeds and paired statistical tests, such as bootstrap testing or paired t-tests, to better assess the robustness of AUC and MRR improvements.
The current system focuses on top–bottom pairing and does not fully support complete outfits that include shoes, bags, accessories, or other fashion items. Mathematically, the framework could be extended by representing a full outfit as a set of items and aggregating compatibility scores across multiple item pairs. For an outfit containing
where
From the computational perspective, the proposed framework uses pretrained language representations, which introduce additional preprocessing and memory-management costs. In our implementation, BERT textual features are precomputed and loaded only when needed, reducing GPU memory pressure during training. GPT-4o and LLaMA are also used offline for attribute extraction and similarity estimation, rather than during online recommendation. Nevertheless, this study does not provide a direct efficiency comparison with PAI-BPR in terms of training time, inference latency, or memory usage. A controlled efficiency benchmark under the same hardware and implementation settings remains future work.
Finally, as the amount of product data increases, the extracted attribute vocabulary and similarity table may also need to be updated. Although richer attributes can provide more detailed product representations, they may increase maintenance cost. Future work can explore more stable attribute schemas, vocabulary merging, and incremental update strategies.
This study proposed an LLM-enhanced multimodal framework for personalized fashion recommendation. The framework integrates visual features, contextual textual representations, and structured attribute features to improve outfit compatibility modeling and user preference prediction. Specifically, a pretrained Japanese BERT encoder was used to replace the conventional Word2Vec- and CNN-based text representation pipeline, while large language models were employed to extract fine-grained fashion attributes and estimate semantic similarities among attribute values.
Experiments on the IQON3000 dataset show that the proposed method outperforms existing baselines. The best-performing configuration achieved an AUC of 0.8477, improving upon the original PAI-BPR baseline. Ranking-based evaluations using MRR and MRR@K further indicate that the proposed model is more effective at placing compatible items near the top of the recommendation list. These results suggest that LLM-based semantic representation and attribute-level modeling provide useful complementary information for multimodal fashion recommendation.
Several limitations remain. The current framework focuses on top–bottom pairing and has not yet been extended to full outfit recommendation involving multiple item categories. In addition, LLM-based attribute extraction and similarity estimation introduce additional computational cost and may depend on the consistency of model-generated outputs. Future work will therefore investigate multi-item outfit recommendation, more efficient attribute modeling strategies, and interpretable mechanisms for explaining recommendation results.
Acknowledgement: The authors gratefully acknowledge Chang Gung University, Taiwan, for providing the GPU computing facilities and hardware resources that made this research possible.
Funding Statement: This work was supported by the National Science and Technology Council (NSTC), Taiwan, under Grant No. 114-2221-E-182-041-MY3 and the partial financial support from Chang Gung Memorial Hospital, Linkou, under Grant No. NERPD4Q0022.
Author Contributions: Ti-Lun Miao: conceptualization, methodology, software, validation, formal analysis, investigation, data curation, visualization, writing—original draft. Hsien-Tsung Chang: conceptualization, methodology, supervision, project administration, funding acquisition, resources, writing—review & editing. All authors reviewed and approved the final version of the manuscript.
Availability of Data and Materials: The dataset used in this study (IQON3000) is publicly available and can be accessed from the source cited in [25]. The dataset are available on GitHub at: https://github.com/hanxjing/GP-BPR.
Ethics Approval: This study did not involve human participants, animals, or identifiable personal data. Therefore, ethical approval was not required.
Conflicts of Interest: The authors declare no conflicts of interest.
References
1. Berti L, Giorgi F, Kasneci G. Emergent abilities in large language models: a survey. arXiv:2503.05788. 2025. [Google Scholar]
2. Elman JL. Finding structure in time. Cogn Sci. 1990;14(2):179–211. doi:10.1207/s15516709cog1402_1. [Google Scholar] [CrossRef]
3. Kim Y. Convolutional neural networks for sentence classification. In: Moschitti A, Pang B, Daelemans W, editors. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); 2014 Oct 25–29; Doha, Qatar. p. 1746–51. [Google Scholar]
4. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4–9; Long Beach, CA, USA. p. 5998–6008. [Google Scholar]
5. Mikolov T, Chen K, Corrado GS, Dean J. Efficient estimation of word representations in vector space. In: Proceedings of the 2013 International Conference on Learning Representations; 2013 May 2–4; Scottsdale, AZ, USA. [Google Scholar]
6. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; 2019 Jun 2–7; Minneapolis, MN, USA. p. 4171–86. [Google Scholar]
7. Radford A, Narasimhan K, Salimans T, Sutskever I. Improving language understanding by generative pre-training. San Francisco, CA, USA; 2018. OpenAI technical report. [cited 2026 Aug 11]. Available from: https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf. [Google Scholar]
8. Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA: open and efficient foundation language models. arXiv:2302.13971. 2023. [Google Scholar]
9. He R, McAuley J. VBPR: visual bayesian personalized ranking from implicit feedback. In: Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence; 2016 Feb 12–17; Phoenix, AZ, USA. p. 144–50. [Google Scholar]
10. Zadeh A, Chen M, Poria S, Cambria E, Morency LP. Tensor fusion network for multimodal sentiment analysis. In: Palmer M, Hwa R, Riedel S, editors. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing; 2017 Sep 7–11; Copenhagen, Denmark. p. 1103–14. [Google Scholar]
11. Liu Z, Shen Y, Lakshminarasimhan VB, Liang PP, Bagher Zadeh A, Morency LP. Efficient low-rank multimodal fusion with modality-specific factors. In: Gurevych I, Miyao Y, editors. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2018 Jul 15–20; Melbourne, Australia. p. 2247–56. [Google Scholar]
12. Nagrani A, Yang S, Arnab A, Jansen A, Schmid C, Sun C. Attention bottlenecks for multimodal fusion. In: Ranzato M, Beygelzimer A, Dauphin Y, Liang PS, Vaughan JW, editors. Advances in Neural Information Processing Systems. Vol. 34, Red Hook, NY, USA: Curran Associates, Inc.; 2021. p. 14200–13. [Google Scholar]
13. Sagar D, Garg J, Kansal P, Bhalla S, Shah RR, Yu Y. PAI-BPR: personalized outfit recommendation scheme with attribute-wise interpretability. In: Proceedings of the 2020 IEEE Sixth International Conference on Multimedia Big Data (BigMM); 2020 Sep 24–26; New Delhi, India: IEEE; 2020. p. 221–30. doi:10.1109/bigmm50055.2020.00039. [Google Scholar] [CrossRef]
14. Alshareet O, Ben Hamza A. Adaptive spectral graph wavelets for collaborative filtering. Pattern Anal Appl. 2024;27(1):10. doi:10.1007/s10044-024-01214-x. [Google Scholar] [CrossRef]
15. Vasileva MI, Plummer BA, Dusad K, Rajpal S, Kumar R, Forsyth D. Learning type-aware embeddings for fashion compatibility. In: computer vision–ECCV 2018. Cham, Switzerland: Springer International Publishing; 2018. p. 405–21. doi:10.1007/978-3-030-01270-0_24. [Google Scholar] [CrossRef]
16. Verma D, Gulati K, Shah RR. Addressing the cold-start problem in outfit recommendation using visual preference modelling. In: Proceedings of the 2020 IEEE Sixth International Conference on Multimedia Big Data (BigMM); 2020 Sep 24–26; New Delhi, India. p. 251–6. doi:10.1109/bigmm50055.2020.00043. [Google Scholar] [CrossRef]
17. Salton G, McGill MJ. Introduction to modern information retrieval. Columbus, OH, USA: McGraw-Hill; 1983. [Google Scholar]
18. Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, et al. Deep contextualized word representations. In: Walker M, Ji H, Stent A, editors. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers); 2018 Jun 1–6; New Orleans, LA, USA. p. 2227–37. [Google Scholar]
19. Yu P, Xu Z, Wang J, Xu X. The application of large language models in recommendation systems. Proc SPIE. 2025;13635(4):103. doi:10.1117/12.3062544. [Google Scholar] [CrossRef]
20. Xu W, Wu Q, Liang Z, Han J, Ning X, Shi Y, et al. SLMRec: distilling large language models into small for sequential recommendation. arXiv:2405.17890. 2024. [Google Scholar]
21. Zhang D, Yu Y, Dong J, Li C, Su D, Chu C, et al. MM-LLMs: recent advances in multimodal large language models. In: Ku LW, Martins A, Srikumar V, editors. Findings of the association for computational linguistics: ACL 2024. Bangkok, Thailand: Association for Computational Linguistics; 2024. p. 12401–30. [Google Scholar]
22. Liu Q, Zhao X, Wang Y, Wang Y, Zhang Z, Sun Y, et al. Large language model enhanced recommender systems: a survey. arXiv:2412.13432. 2024. [Google Scholar]
23. Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, et al. Measuring massive multitask language understanding. In: Proceedings of the 2021 International Conference on Learning Representations; 2021 May 4; Vienna, Austria. [Google Scholar]
24. OpenAI. GPT-4 technical report. arXiv:2303.08774. 2023. [Google Scholar]
25. Song X, Han X, Li Y, Chen J, Xu XS, Nie L. GP-BPR: personalized compatibility modeling for clothing matching. In: Proceedings of the 27th ACM International Conference on Multimedia 2019 Oct 21–25; Nice, France. p. 320–8. doi:10.1145/3343031.3350956. [Google Scholar] [CrossRef]
26. OpenAI. GPT-4o system card. arXiv:2410.21276. 2024. [Google Scholar]
27. Meta. LLaMA 3.3 70B Instruct; 2024. Instruction-tuned 70B multilingual LLM released December 6, 2024 [cited 2026 Aug 11]. Available from: https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct. [Google Scholar]
28. Tohoku NLP Group. tohoku-nlp/bert-base-japanese model; 2023. Pretrained BERT base model for Japanese, Tohoku University NLP Group. [cited 2026 Aug 11]. Available from: https://huggingface.co/tohoku-nlp/bert-base-japanese. [Google Scholar]
29. Rendle S, Freudenthaler C, Gantner Z, Schmidt-Thieme L. BPR: Bayesian personalized ranking from implicit feedback. arXiv:1205.2618. 2012. [Google Scholar]
30. Han X, Wu Z, Jiang YG, Davis LS. Learning fashion compatibility with bidirectional LSTMs. In: Proceedings of the 25th ACM International Conference on Multimedia; 2017 Oct 23–27; Mountain View, CA, USA. p. 1078–86. doi:10.1145/3123266.3123394. [Google Scholar] [CrossRef]
31. Song X, Feng F, Liu J, Li Z, Nie L, Ma J. NeuroStylist: neural compatibility modeling for clothing matching. In: Proceedings of the 25th ACM International Conference on Multimedia; 2017 Oct 23–27; Mountain View, CA, USA. p. 753–61. doi:10.1145/3123266.3123314. [Google Scholar] [CrossRef]
32. Cao D, Nie L, He X, Wei X, Zhu S, Chua TS. Embedding factorization models for jointly recommending items and user generated lists. In: Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval; 2017 Aug 7–11; Shinjuku, Tokyo, Japan. p. 585–94. doi:10.1145/3077136.3080779. [Google Scholar] [CrossRef]
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools