Please use this identifier to cite or link to this item: https://bura.brunel.ac.uk/handle/2438/33697
Title: Intelligent outlier detection under imperfect industrial data conditions
Authors: Fang, Jingzhong
Advisors: Wang, Z
Liu, X
Keywords: Intelligent Data Analysis;Data Quality;Artificial Intelligence;Deep Learning;Weakly Supervised Learning
Issue Date: 2026
Publisher: Brunel University London
Abstract: With the growing complexity and data intensity of modern industrial systems, intelligent data analysis (IDA) has become essential for reliable data interpretation and efficient operation. By leveraging advanced analytical and computational techniques, IDA can handle complex, largescale data sets, thereby providing valuable insights for industrial operations and significantly improving the stability and efficiency of industrial processes. Nevertheless, data quality remains a fundamental prerequisite for IDA performance, as it directly affects the accuracy and trustworthiness of derived insights. Due to complex operating conditions, high cost, and limited availability of expert annotation, high-quality data is often difficult to acquire. The imperfections in the data can affect training and evaluation, thereby reducing the overall performance in real-world applications. In practice, imperfections in industrial data generally fall into two main categories: 1) data scarcity, such as class imbalance or limited sample size; and 2) label quality issues, such as missing labels or noisy labels. In this thesis, we deal with the above-mentioned data imperfections arising from the data acquisition process and the annotation process to develop innovative approaches for robust and reliable IDA in industrial applications. It should be pointed out that all approaches developed in this thesis have been evaluated and applied to outlier detection on imperfect industrial data collected from real-world wire arc additive manufacturing (WAAM) processes. • To achieve outlier detection on unlabeled data, an improved optimization-based clustering algorithm is proposed, where the initial locations of the cluster centroids in the fuzzy C-means algorithm are optimized by an improved particle swarm optimization (PSO) algorithm. An adaptive switching randomly perturbed particle swarm optimization (ASRPPSO) algorithm is developed to enhance the convergence rate and particle’s search ability of the PSO algorithm, where a distance-based weighting strategy and switching strategy are introduced to update parameters of the optimizer. The particles can conduct a thorough search by accounting for both evolutionary states and the distances to the global and personal best positions, thereby improving the convergence rate and solution accuracy. Via the ASRPPSO-based selection of optimal clustering centroids, the proposed algorithm does not rely on centroid initialization, thereby facilitating a better cluster partition. • With the aim to guarantee the outlier detection performance on imbalanced data under the small sample problem, an optimized deep transfer learning framework is developed, where a novel deep domain adaptation strategy is designed to minimize the cross-domain discrepancies, the weighting factors are designed to handle the data imbalance problem, and the PSO algorithm is utilized for hyper-parameters tuning. By leveraging the domain knowledge and optimal hyper-parameter selection, the developed framework effectively balances performance and efficiency when dealing with outlier detection tasks on imbalanced data under the small sample problem. • To handle the noisy label problem under limited-data conditions in outlier detection, a novel Transformer-embedded learning with noisy labels framework with fuzzyclustering- assisted contrastive learning (TFCCL) is developed, where a fuzzy-clusteringassisted contrastive learning approach, a dynamic two-stage training scheme and a joint learning strategy are introduced to train the outlier detector. The TFCCL framework integrates supervised learning with contrastive learning, thereby reducing reliance on potentially noisy labels and enhancing model robustness. • For the purpose of outlier detection under the noisy label problem when sufficient data are available, a role-differentiated learning with noisy labels (RD-LNL) approach is put forward, where a leader-follower-inspired sample selection (LFSS) strategy is proposed for identifying potential clean samples for model training. A selection metric and an adaptive selection scheme are designed to adjust the number of selected samples. By combining the joint learning and adaptive sample selection, the RD-LNL framework achieves robust outlier detection in the presence of noisy labels. • To fully evaluate the application potential of the developed frameworks. All the developed frameworks are applied to outlier detection tasks on WAAM data sets collected from real-world manufacturing processes to examine their adaptability, robustness, and generalization capability. The experimental results demonstrate the effective and robust performance of the developed outlier detection frameworks on WAAM data sets across various data imperfection conditions.
Description: This thesis was submitted for the award of Doctor of Philosophy and was awarded by Brunel University London
URI: https://bura.brunel.ac.uk/handle/2438/33697
Appears in Collections:Computer Science
Department of Computer Science Theses

Files in This Item:
File Description SizeFormat 
FulltextThesis.pdf7.18 MBAdobe PDFView/Open


Items in BURA are protected by copyright, with all rights reserved, unless otherwise indicated.