Skip to content

Classification Nodes

Classification nodes assign spectra or feature rows to classes.

Nodes

Node Use When Inputs Outputs Key Configuration
Train KNN Classifier (classification.knn) You need a distance-based classifier after appropriate scaling or dimensionality reduction. Validation is owned by an explicit split plan outside this fit-only node. X: Array2D; y: Categorical? fitted_state; calibration predictions; probabilities; metrics; distances; neighbor_indices; train_accuracy; plots n_neighbors; weights; metric; scale.
Apply Fitted KNN (classification.apply_knn) Apply the exact KNN reference set, scaling, distance, and voting state to new rows. X_new: Array2D; fitted_state: ClassificationModel y_pred: Categorical; y_prob: Array2D no parameters. Accepts only the KNN serializer.
Train PLS-DA Classifier (classification.plsda) You want a latent-variable classifier for exactly-one-class spectral discrimination. Validation is owned by an explicit split plan outside this fit-only node. X: Array2D; y: Categorical? fitted_state; calibration predictions; raw response class_scores; X_scores; X_loadings; calibration metrics; plots n_components; scale. Uses the native Sherpa SIMPLS authority.
Apply Fitted PLS-DA (classification.apply_plsda) Apply fitted PLS2 coefficients and the declared maximum-response decision rule to new rows. X_new: Array2D; fitted_state: ClassificationModel y_pred: Categorical; class_scores: Array2D no parameters. Scores are responses, not probabilities; accepts only the PLS-DA serializer.
Train SIMCA Classifier (classification.simca) You want class-specific PCA models that can reject samples outside known classes. Validation is owned by an explicit split plan outside this fit-only node. X: Array2D; y: Categorical? fitted_state; calibration predictions; class_assignment; distances; class_distance_matrix; confusion_matrix; plots; metrics n_components; confidence_level; critical_limits_method. Uses the native Sherpa SIMCA authority.
Apply Fitted SIMCA (classification.apply_simca) Apply explicit class PCA models and their complete T²/Q limits to new rows. X_new: Array2D; fitted_state: ClassificationModel y_pred: Categorical; class_affinity: Array2D no parameters. Can return unassigned; affinity is not probability; accepts only the SIMCA serializer.

Key Outputs

Expect predictions, class probabilities, response scores, or affinities where appropriate, plus confusion matrices and split-aware metrics when validation is part of the workflow. Application is deliberately family-specific: no node may guess the model family, invent missing SIMCA limits, or reinterpret a numerical output as a probability.

KNN fits the exact configured neighbor count, distance rule, weighting rule, and optional autoscaling state. The node reports calibration-fit diagnostics, not validation performance. Use an explicit stratified or grouped split plan and the shared fold executor for validation; preprocessing and the classifier are then fit only on each training partition. The saved artifact carries the reference rows, encoded class identities, and fitted scaling vectors needed to reproduce scikit-learn's exact zero-distance and tie behavior. Scientific reference: Cover and Hart, Nearest Neighbor Pattern Classification, IEEE Transactions on Information Theory 13 (1967), 21–27, DOI 10.1109/TIT.1967.1053964.

PLS-DA fits PLS2 to a 0/1 dummy class table and assigns each sample to the largest predicted dummy response, following the usual rchemo::plsrda exactly-one-class rule. Its class_scores are regression responses, not posterior probabilities. The node performs one fit and carries no hidden cross-validation setting. The standard PLS-DA analysis template uses an explicit stratified train/test split; repeated or grouped validation uses the shared split-plan and fold-executor workflow. This applies to every dataset, not only grouped instrument studies. Report the class order, split record, confusion matrices, and fold-local metrics. Scientific reference: Barker and Rayens, Partial least squares for discrimination, Journal of Chemometrics 17 (2003), 166–173, DOI 10.1002/cem.785.

SIMCA Note

SIMCA is not just nearest-class classification. It can reject samples that do not fit any class model, which is important for QC workflows. Its training node performs one fit and reports calibration diagnostics only. Validation belongs to an explicit split plan and the shared fold executor; an unassigned result is a scientific rejection, not a fabricated class prediction.