SpectraSherpa 0.6.0
0.6.0 release lifecycle. Install 0.6.0 from PyPI only after the public index reports that exact version. Before the public tag exists, use only the exact monorepo commit named by the qualification record. After the tag exists but before PyPI reports 0.6.0, use the exact
spectra-sherpa-v0.6.0source tag. Source version text alone is not publication evidence.
After the public index reports 0.6.0, install the complete native Workbench and Python SDK together:
python -m pip install "spectra-sherpa==0.6.0"
spectra-sherpa
For a pre-publication qualification candidate, use the exact monorepo commit named by the qualification record. From that already-pinned checkout:
cd packages/spectra-sherpa
python -m pip install poetry
poetry env use python3.11
poetry install
npm --prefix frontend ci # requires Node.js 22
npm --prefix frontend run build
poetry run spectra-sherpa
The Workbench opens locally, normally at http://127.0.0.1:8000. The same
environment immediately supports canonical Python workflows:
poetry run python -m spectra_sherpa.examples.canonical_workflow
Add only the capabilities you need
poetry install --extras "scp"
poetry install --extras "hitran"
poetry install --extras "nist"
poetry install --extras "scp hitran nist"
For a qualification candidate, these commands exercise source rather than the public wheel/source distributions, checksums, and PyPI attestations. Signed/notarized desktop installers ship with 0.6.0 for Windows and macOS. The desktop, commercial, and trial paths share the 0.6.0 scientific core; Hybrid is separately distributed and its clients and UI are absent from OSS. Each remains subject to its own signing, deployment, and acceptance gates.
Managed trial boundary
The hosted demo profile is a supplied-public-data trial.
It starts after six-digit email verification, ends on day 31 regardless of
activity, and may delete work after seven consecutive inactive days. It
accepts no payment method, exposes no checkout or automatic renewal, and
accepts no general customer scientific uploads or imports. Its one ingress
exception is the dedicated registered-reference selector: the scientist
downloads a qualified artifact directly from its provider, and Sherpa admits
it only when exact size, SHA-256, native structure, and scientific projection
match the active registry. This path-specific classification does not constrain
ordinary supported-file ingestion in local OSS or the separate Pro profile.
Managed Generative Advisor, both governed campaign lanes, and Campaign Review
Packages remain available within trial quotas. Contact
support@spectrascientific.ai after expiry.
Existing registered-reference imports whose metadata predates source_scope
remain readable when every other field matches the current registry. Sherpa
derives the missing description from that registry without rewriting the stored
sidecar; source-byte and scientific-identity validation remain unchanged. This
specific compatibility case does not require deleting and re-importing the
dataset. Any other metadata difference still requires investigation.
The same 0.6.0 codebase supports a separate Pro Standard profile with customer uploads, managed Advisor, workflow proposals, eligible optimization campaigns, and signed application delivery for local inference. Its availability depends on its own payment and deployment qualification; demo restrictions do not define Pro's upload or optimization entitlement. Private Team/Hybrid delivery is outside the OSS release and requires separate qualification.
Local OSS provides user-configured BYOK chat and the scientific workbench.
Managed project-aware Advisor and campaign execution are hosted capabilities.
Optimize offers signed package and publisher-key downloads, links to the
source run/workflow, and opening the selected result as a separate workflow.
A candidate opened in Workflow keeps the campaign's validation scope: its
evaluator is scored by cross-validation on the campaign's exact folds, never on
the rows the model was fitted on. If the sheet's data no longer yields those
folds, the run stops instead of scoring different partitions.
Optimize's verification report lists four outcomes: package integrity and the publisher
signature are checked when the report is downloaded; validation and application
are recorded results (with the exact dataset and split digests) marked not_run,
because they are reproduced only by the recipient's local import.
Reference-file acquisition and portability are described in Datasets and Providers. For the boundaries between current-graph export, historical evidence, full-project objects, and signed application packages, see Export Design.
Validation and reporting additions
The final v0.6.0 validation baseline carries the full review evidence through the Workbench and exports:
- nested selection and model fitting remain fold-local, with explicit group authority available at both outer and inner levels;
- repeated validation retains repeat, fold, seed, prediction, selected component and selector-stability evidence, while reporting spread as partition sensitivity for one dataset rather than as extra specimens;
- recorded specimen/batch identity comparisons are labelled separately from operator declarations and unknown evidence, and do not prove independent sampling by themselves;
- saved-run HTML, Markdown and JSON reports share the same metric formulas, target units, population, exclusions and claim limits; and
- row-level reference/prediction/residual values and figures are explicit opt-in outputs from retained evidence. Figures decline above 5,000 rows without silent subsampling, while numeric evidence remains available.
The same boundaries apply to campaign verification and analytical qualification. Package integrity, publisher authentication, validation reproduction, fitted-application reproduction and a human qualification decision remain separate outcomes. See Reports and Exports and Validation and Model Application for the user-facing workflow.
Compatibility disclosures
- Balanced accuracy now averages recall only across true classes observed in the evaluated rows. Predicted-only rejection or otherwise unobserved labels still count against ordinary accuracy and the applicable true-class recall, but an empty true-class row no longer dilutes the balanced mean. Values from before and after this change may therefore not be directly comparable when rejection or unobserved labels are present.
-
The external
POST /workflows/{id}/predictendpoint has been removed. Use the durable model-application replacements underPOST /api/v1/runs/batch*:/api/v1/runs/batch,/api/v1/runs/batch/folder, or/api/v1/runs/batch/files/{artifact_uid}. -
Native PCA, rubberband, PLS, and PLS-DA ship in the default installation. The optional
scpextra adds EFA, MCR-ALS, and SIMPLISMA. Vendor-file ingestion is admitted only after the corresponding native reader is qualified; installingscpdoes not add a hidden ingestion fallback. - Train/Test Split now consumes an exact attached validation-group authority. Random, stratified, and sequential designs hold complete groups out and bind their identities in the split receipt, preventing specimen, batch, or block replicates from crossing train and test. Observation-level Kennard–Stone, DUPLEX, and SPXY explicitly refuse grouped input rather than inventing an unqualified group-distance method. The Workbench names the active grouping, disables those incompatible choices before execution, and shows the exact held-out groups from the bound split plan afterward.
- A scientist's chosen response is now one inseparable
TargetAuthority: column identity, continuous/categorical type, units, and the exact source or collection SHA-256 travel together from My Dataset through template admission, saved DAGs, execution provenance, canonical campaign baselines, and executable/notebook exports. A changed source, mismatched units or type, stale cross-dataset selection, or separately edited field is refused before analysis. Pre-0.6 saved graphs that persisted the former independentselected_targetandtarget_typeparameters must be reopened and rebound; those loose fields are intentionally not migrated into scientific authority. - Native ingestion now attaches one parser authority at the format-registry boundary, so previews, workflow loading, SDK calls, and durable exports carry the same format, variant, parser version, source-member digests, asset identity, and warnings. Every admitted asset must also pass the same closed conformance checklist for identity, numeric shape, dimension roles, canonical feature-axis semantics, predictor role, and disclosed missingness; a reader cannot return a partially described dataset to any consumer.
- Cross-validation and held-out regression evaluations now use the same
version-2 metric registry. Both define bias as prediction minus observation
and RER as observed reference range divided by bias-corrected SEP; both expose
SEP, slope, and intercept under the same definitions. The out-of-fold report
advances to
spectrasherpa-cross-validation-evaluation/3. - Native PCA uses deterministic full-SVD fitting. Large datasets that previously selected a randomized SpectroChemPy solver can therefore produce different fitted scores and loadings from pre-0.6 runs.
- Saved ordinary PLS-DA runs with exact retained evidence can enter a managed candidate campaign and settle under the classification objective. A refusal now identifies incomplete retained evidence, an inadmissible graph shape, or a mismatched pinned numerical runtime separately; the runtime requirement is not weakened to admit a different local package version.
- Sherpa can now turn a scientist's bounded verbal optimization direction into a small, review-only set of structural workflow hypotheses. The managed LLM sees aliased node references, editable settings, bounded axis/target/count metadata, the objective, and the live managed-operation vocabulary—not raw spectra, row labels, sample identities, paths, credentials, internal target names, or fitted state. The server deterministically compiles each mutation, revalidates node settings and task semantics, re-admits the typed workflow, removes duplicate directions, and computes the displayed seed-relative diff and resource envelope. A proposal cannot admit a campaign, open a dataset capability, or execute work; the parameter-only proposal remains available for the later factorial stage. Every structural hypothesis now freezes its server-derived seed comparator, paired-fold/group calculation, meaningful improvement threshold, guardrails, stability rule, and evidence/failure dispositions before execution. The unchanged seed is included and budgeted as the reference trial. Managed text passes a centralized fail-closed egress policy, graph deduplication ignores incidental node IDs, interval bounds are checked on the final graph, and the live catalog exposes only settings the managed validation authority will actually admit. Structural planning now starts from an active campaign ID, selected seed ID, and exact development- examination ID rather than caller-supplied run, workflow, or objective combinations. The server reproduces the complete examination against the campaign's qualified Rev0 custody before contacting the LLM and persists that full lineage with the proposal. Expected evidence is explicitly limited to tentative development cross-validation hypotheses; explicit unsupported factorial attribution and enumerated assertive confirmation, held-out, development-result, and guaranteed-improvement claim forms are refused. Text screening is bounded; scientists must still review generated hypotheses rather than treating a passed proposal as measured evidence. The egress rules carry a versioned digest, and internal compiler faults stop for review without being retried or recorded as invalid provider output. The saved-run campaign page now turns this authority into a conversational planning workspace: it keeps one digest-bound current valid preview, retains immutable parented revisions, sends only a bounded data-free summary of that preview back to Sherpa, and lets the scientist reshape directions, compare alternatives, cap latent variables, or reduce total trials. Rejected or stale edits leave the prior plan current. Hypothesis rationale, expected development evidence, trade-offs, topology-complete workflow changes, coverage, and estimated compute remain visible. Cost is shown when backed by server pricing; otherwise the page says that cost is unavailable rather than treating a null estimate as zero. Planning still creates no execution batch and spends no compute. After that review, one explicit Execute action now freezes the exact proposal and a short-lived server-issued resource quote. The quote distinguishes trial count, aggregate compute, orchestration wall deadline, concurrency, estimated cost, personal allowance, deployment allowance, and unavailable confirmation consultation accounting. The server atomically records the scientist's versioned approval receipt, both budget scopes, the immutable execution batch, its candidates, and the durable worker job; a failed transaction leaves none of them behind and an idempotent replay cannot duplicate them. Real terminal results settle the reviewed operation and budget record before promotion.
- Managed optimization in 0.6.0 ends with a scientist-selected final
development checkpoint, an all-development-data application refit, and one
explicitly consumed consultation against independently held confirmation
data. The page shows the frozen workflow and protocol identities, consultation
count, compute allowance, estimated infrastructure cost, and
$0billed amount before access; only aggregate terminal evidence returns. Once that result is revealed, later planning revisions and their descendant iterations and seeds are visibly confirmation-exposed and cannot support a fresh confirmation claim without new independent data. Factorial search remains disabled until its qualification gate closes: no 0.6.0 proposal admits a factorial stage or receives a design matrix, and unsupported main-effect, factor-interaction, isolated-factor causality, best-level, and top-decile attribution is refused while legitimate scatter-effect and matrix-effect language remains available. A qualifying infrastructure failure that reveals no scientific result does not reset confirmation history and cannot be worked around with a fresh campaign. An organization owner or admin may authorize one receipt-bound exact retry of the unchanged frozen request; scientific failures, changed requests, and any post-disclosure retry remain refused. - Renishaw WiRE maps load as one sample-major matrix of shape
(spectra, points), not the(rows, columns, points)cube SpectroChemPy returns. Every spectrum stays one chemometric sample and WMAP topology is explicit metadata, so the grid is reconstructible; workflows that indexed the third dimension must be updated. - Vendor feature axes are rebuilt in float64 from each file's stored float32 endpoints and point count. Axis values therefore differ from pre-0.6 runs at float32 resolution -- under 1e-3 cm-1 on the qualified corpus. Intensities are unchanged and remain bit-exact against the readers being replaced.
- OMNIC SPG groups order samples by their explicit acquisition timestamp and retain the original directory index as sample metadata. Files whose directory order differs from acquisition order will present rows in a different order than pre-0.6 runs.
- Qualified legacy OMNIC SPA now retains bounded operator comments, processing
history, experiment/accessory information, and qualified embedded-source
identity when present. Missing SPA values are admitted only when the retained
code-23 profile's NaN mask exactly matches an explicit OMNIC processing-history
blank interval; arbitrary NaNs, mismatched masks, and infinities refuse. The
affected Nicolet Apex files retain both the instrument-displayed
Transmittancequantity and the separateFinal format: Single Beamhistory, but ratio units remain unresolved, so transmittance-to-absorbance conversion refuses. Numerical preprocessing also refuses missing input; usepreprocess.clip_rangeto retain the measured region rather than imputing it. Same-source SPA/CSV/JDX/SPC parity for these files is claimed only over the common measured finite region. - CSV import now has four durable supplier-neutral profiles: headered or unheaded coordinate/intensity, each with decimal-point or decimal-comma values and comma/semicolon/tab delimiter support. Ambiguous unheaded X/Y files require explicit confirmation, and that exact choice survives upload, workflow execution, and SDK export.
- An OPUS file may contain a data block SpectraSherpa recognizes structurally
but has not independently qualified. Such a block is refused individually,
named in
opus.refused_blockswith its offset, type code, and reason, and reported as an ingestion warning; the qualified blocks beside it still load. A file whose blocks are all unqualified is still refused entirely. - An OPUS data block whose pairing key uniquely identifies one status block is
admitted even if that status block's recorded MNY/MXY does not reproduce
from the block's decoded bytes -- a real, published acquisition exhibits
this. A data block sharing its pairing key with more than one qualified
status block -- a repeated result type, which real acquisitions do store --
is now disambiguated by decoded value instead of being refused outright.
When more than one same-type result is admitted from one file, the
last-written occurrence receives the bare asset id (matching Bruker's own
convention, "last spec is OPUS preference") and earlier ones are suffixed
(
a,a_2, ...). No pre-0.6 run admitted such a file at all -- every same-type duplicate was refused outright -- so this is new coverage, not a changed result. hitranadds live HITRAN/HAPI acquisition.nistadds live NIST WebBook acquisition.- Fourteen non-imaging projections across five Eigenvector Research, Inc. teaching-data families are visible as five user-acquired dataset packages: CGL NIR; Corn with M5/MP5/MP6 Data Views; Diesel D4052; Metal Etch with machine-sensor/OES/RF-monitor Data Views; and NIR Shootout 2002 with calibration/test/validation Data Views for both instruments. Each package card gives one technical sentence and one import action. One provider-level Browse all Eigenvector datasets link opens the complete Eigenvector catalog; individual cards do not link to provider files. Art Image Data A and IASIM16 identities remain in the qualification registry but are not listed because 0.6.0 does not offer a guided hyperspectral-image workflow or analysis starter. An expert local workflow may explicitly select a qualified DSO image-cube result and manually add the native masked PARAFAC node; the exact registered IASIM16 Test 1 source was privately qualified on that path. Sherpa does not fetch, proxy, mirror, package, or redistribute provider archives.
These references are not pip extras or Sherpa-hosted data. One Corn archive
materializes all three aligned instrument views under one package, preserves
one shared 80-specimen identity and the complete Moisture, Oil, Protein, and
Starch annotation table, and initially checks only M5 for display. One NIR
Shootout archive similarly materializes all six instrument/cohort views with
Weight, Hardness, and Assay. Metal Etch requires its three independently
downloaded archives and keeps their incompatible feature spaces separate;
cross-view row alignment is not inferred. No property is silently selected
as the analysis target; the scientist makes that choice in My Dataset.
The SpectroChemPy extra does not provide a product data catalog. Campaign
Review reproduction can use the exact provider-downloaded artifact plus its
qualified projection ID; no Spectra-supplied Corn fixture is required.
- model.parafac is a canonical, deterministic, local-only DAG node for
explicitly modeled rank-3 through rank-6 data. It has no template reference in
0.6.0. The qualified three-mode HSI path preserves both spatial modes, honors
source-declared spatial exclusions during every ALS update and reconstruction
diagnostic, and records unfolding: none. It is exploratory decomposition,
not a target-detection claim.
Scientific-data correction: pre-release Eigenvector CSV labels
Two Eigenvector teaching datasets were mislabeled in early, pre-release legacy CSV materializations. This correction changes property meaning, not merely display text:
- CGL NIR:
Dry Gluten, Moisture, Protein, Wet Glutenwas wrong. The authoritative properties areCasein, Glucose, Lactate, Moisture. - NIR Shootout 2002 (calibration, test, and validation views for both
instruments):
Active, Hardness, Weightwas wrong. The authoritative properties areWeight, Hardness, Assay.
No released user database exists, so there is no released-data migration. Any saved pre-release dataset, workflow, model, or export built from a manually retained CSV with either old exact header set remains scientifically mislabeled; re-import the provider source, verify the corrected headers, and rebind or rerun that work before interpreting results. SpectraSherpa does not silently rewrite raw project files or existing fitted artifacts.
Previously acquired content-addressed reference data remains usable without its acquisition client. Missing optional capability fails before scientific execution and names the exact install command; no fallback algorithm runs.
Dependency security
The 0.6.0 lock authorities include patched releases for the default Python
installation, optional nist, scp, and hitran trees, build dependencies,
and the complete frontend toolchain. The exact SpectroChemPy 0.8.1 scientific
profile remains unchanged. Release CI blocks on live vulnerability findings
from both the all-extras Python resolution and the complete npm lockfile; the
runtime-only subset is not used to waive vulnerable optional or build tools.
What is new
-
Data → Multi-well now owns complete multi-well experiment intent: samples, mixture recipes, experimental factors, plate layout, acquisition order, matching rules/results, and a final review. Settings provides the user-owned default plate format. My Dataset projects saved acquisition matches into a proposed measured-sample table, shows exact before/after differences, and requires a separate Publish new immutable version action. Both pages report the same synchronization state. Legacy DOE mixtures/components, factors, run levels, matches, and user presets migrate into this storage before their DOE-only ORM/service layer is retired. The separate mutable sample catalog is deliberately retained as the authority for specimen metadata; plans snapshot only referenced specimen identities with stable legacy references and do not mutate catalog rows.
-
One canonical registry and executor now serve Workbench, Python, managed workers, export, import, and reproduction.
- The workflow-template catalog preserves every distinct v0.5.30 starter
purpose while removing three misleading duplicates.
pca_exploratoryandspectral_decompositionare consolidated into PCA Exploration & Outlier Diagnostics, which adds three-component scores, loadings, explained variance, and T-squared/Q-residual review. The WIPir_opus_analysiscanvas is retired because OPUS is a native input format rather than a scientific method; its range clipping, baseline, normalization, peak detection, and summary operations remain available as ordinary nodes and starters. The 0.6 catalog retains 23 v0.5.30 slugs and adds 10 new starters, for 33 total; none of the three retired slugs represented a unique analysis algorithm. - My Dataset can overlay multiple checked packaged datasets without
combining or rewriting their scientific records. Clicking a package name
still focuses its files, acquisition settings, labels, metadata, quality,
and workflow context; the checkboxes control only the shared display. Each
focused package also exposes per-file plot checkboxes, while its package
checkbox shows checked, clear, or partial state. The plot sits immediately
below both selection lists; the list headings stay fixed while their rows
scroll. Hover replaces the incomplete legend and carries both the existing
sample identity and the exact source filename for every spectrum. The one
file focused for inspection is emphasized in the plot without changing its
plot-inclusion state. File-list hover shows the corresponding sample label
without requiring a file to become active.
The view preserves exact feature axes, colors common
specimen_idvalues consistently, refuses incompatible declared quantities or units, and remains bounded to 12 packages and 50 visible spectra. Global display-metadata controls sit directly below the plot, before the five focused detail groups collapsed under an explicit active-dataset heading. - The base install reads qualified Bruker OPUS, Galactic SPC, legacy Thermo OMNIC SPA/SPG/SRS, and Renishaw WiRE WDF sources through bounded native readers. Multi-result files preserve explicit asset identity, and WDF maps retain exact sample coordinates rather than being flattened into an invented image grid.
- The supported Python SDK covers datasets, preprocessing, modeling, selection, validation, artifact application, workflow execution, plot specifications, and offline project/evidence inspection.
- Canonical project archives can be inspected, compared, and re-exported offline. Integrity, publisher authentication, validation reproduction, and application reproduction are reported separately.
- Completed hosted optimization campaigns are downloadable as one signed, data-free Campaign Review Package. The package contains the complete frozen candidate ledger, explicit failed/refused candidates, selection decision, terminal full-data refit, fixed application DAG/artifact, and publisher attestation. The OSS CLI and Workbench inspect the package without an account, reproduce the recorded held-out evidence, and apply the one fixed artifact; they cannot create, search, select, settle, resume, or rerun a managed campaign.
- Governed optimization has two explicit bounded lanes: quantitative calibration (native PLS candidates) and categorical classification (three PLS-DA and three SIMCA configurations). Reference-set classifiers such as KNN are excluded from the data-free Campaign Review deliverable because their fitted state retains calibration rows; KNN remains available for local Workbench use. Validation is deterministic and group-aware when groups are declared. Failed candidates remain visible rather than disappearing from the ledger, and terminal refit remains separate from held-out evidence.
- Optimization studies now retain a long-lived, scientist-owned lineage above each immutable execution batch. Rev0 and every explicitly promoted revision are declarative, data-free workflow seeds linked to their development evidence, parent seed, examination identity, and the scientist's rationale. Iterations can branch from an earlier promising seed. A change to the dataset, response, grouping, split plan, objective, runtime, operation registry, metrics, or evidence schema starts a new examination epoch and resets return comparisons. Full-data refits and fitted estimator state are never admitted as validation seeds. Trial position is now called its ordinal; the former overloaded candidate field is removed.
- The supplied Lavender starter and user-acquired registered references use the same saved-DAG and artifact authorities. Corn has no privileged product admission: like every other Eigenvector reference, it passes the same exact registry, native materialization, grant, export, and deletion path. The release campaign evidence happens to use Corn M5 as its regression exercise; availability of the other registered datasets is not a campaign-performance claim. Dataset cards remain model-neutral: supervised starters visibly name a missing numeric target or categorical class-label requirement and cannot be created until it is satisfied, while target-free data remain available for PCA and other unsupervised work. These compatibility decisions establish structural readiness only: they do not pre-claim adequate sample counts, class balance, fold feasibility, residual-moment sufficiency, or scientific fitness for a particular question. Those checks remain authoritative at workflow execution. Technique matching is advisory and is reported as a recommendation rather than used to admit or refuse a structurally valid dataset. My Dataset shows this same runtime structural profile for built-in, registered-reference, and ordinary user datasets. Hover details name every analysis in each readiness group. Saving and explicitly binding a portable sample-table CSV revalidates its exact source-file and row identities, persists the binding, and immediately recomputes the profile; JSON metadata is not an alternate binding path. The binding can be removed without deleting either file. Incomplete target metadata or one invalid template contract now produces a visible, isolated readiness diagnostic while the successfully parsed data preview remains available. Lavender remains a grouped workflow demonstration—not an authenticity, supplier, population, future-lot, probability, or general-performance claim. Third-party raw bytes are not retrieved, redistributed, or relicensed by SpectraSherpa.
- Native Sherpa SIMPLS is the shared PLS/PLS-DA linear-algebra authority, with multi-metric regression and classification outputs.
- KNN, PLS-DA, and SIMCA training nodes no longer run hidden row-shuffled cross-validation. They report calibration-fit diagnostics only; validation uses the visible split/apply/evaluator graph or the grouped fold executor. Legacy KNN, PLS-DA, and SIMCA nodes are refused: their former CV-derived output ports cannot be translated into current calibration diagnostics without changing scientific meaning. Recreate those nodes and add an explicit validation plan before upgrading. Newly persisted graphs carry an exact validation-semantics authority so future admission is unambiguous.
data.train_test_splitnow consumes the exact bound validation-group authority for random, sequential, and stratified holdouts, records group identities and digests in its split plan, and proves no group crosses the partition. Grouped Kennard–Stone, DUPLEX, and SPXY requests refuse because those row-ranking algorithms do not provide a group-preserving contract.- Regression metric registry version 2 adds bias-corrected SEP, predicted-on- observed slope and intercept, and RER to RMSE, MAE, bias, and R². PLS2 evaluation reports a separate complete metric record for every named response instead of refusing multi-response predictions or blending unlike units. Managed campaign ranking deliberately remains single-target.
- Classification metric registry version 2 adds per-class sensitivity and
specificity derived from the complete confusion matrix. Retained PLS-DA and
KNN workflows may carry a qualified preprocessing chain into leakage-safe
development validation. SIMCA uses a distinct class-modelling objective:
it retains own-class acceptance sensitivity plus unassigned and multiple-
acceptance rates, and reports specificity as unavailable unless a future
examination binds a representative challenge population. Historical saved-
run projection
/1and classification validation/3remain parseable, digest-verified audit authorities. They are explicitly non-executable under profile/8and are never reinterpreted as current results. New qualified records use projection/2and validation/4. - A saved Rev0 holdout score is omitted from Sherpa's optimization seed context when a fitted transform (such as autoscaling, MSC, EMSC, or OSC) learned from all rows before the saved split. Its historical run and result remain inspectable, but that score is not evidence of leakage-safe improvement. New managed trials fit preprocessing inside development folds. Protected one-shot confirmation now supports SIMCA through the frozen fitted artifact and its class-modelling objective: own-class acceptance sensitivity, unassigned rate, and multiple-acceptance rate. It does not substitute ordinary classification accuracy or assert specificity without a representative challenge population. SIMCA's version-2 confirmation bootstrap preserves fitted-class support: ungrouped draws keep each class's observed count, and grouped draws retain only draws covering every fitted class. Thus its unassigned and multiple-acceptance rate intervals describe uncertainty at the observed class mix, not uncertainty from a changing class composition.
- Executable Python and notebook exports now preserve current
data.collection_loadsources. Exact selected ordinary files travel in the export; registered third-party bytes remain external references resolved by exact size and SHA-256. Collection labels, target choice, validation groups, scientific identity, and a path-free native ingestion authority are reconstructed through the same readers and supervision authority used by Workbench. An export refuses if the current reader changes the bound format, variant, parser version, asset identity, source-member digests, warnings, or ingestion-invariant set. A byte-identical provider file may still be renamed and relocated explicitly. - Registered references now distinguish provider-acquired artifact bytes, scientist-facing dataset packages, and target-neutral data views. The Corn package declares M5, MP5, and MP6 as aligned measurements of the same 80 specimens and selects M5 initially for display only; this authority does not choose a target, model, preprocessing operation, or validation method.
- Feature axes now carry a canonical physical quantity separately from their
units. Reader-boundary normalization maps equivalent spellings such as
cm-1,cm⁻¹,cm^-1, and1/cmto canonicalcm-1while preserving the original source spelling for display. Absolute wavenumber and Raman shift remain incompatible quantities despite sharing a unit. Collection assembly, calibration transfer, library comparison, and saved-model application refuse mismatches; compatible grid differences direct the scientist to the explicit, provenance-recordedpreprocess.wavenumber_alignoperation. - Multiway spectral arrays now retain closed dimension roles through ingestion
and analysis. PCA may unfold them only through its explicit, provenance-
recorded
projects_to_2dpolicy. The new nativemodel.parafacoperation instead preserves ranks 3–6, keeps every non-observation mode distinct, reports that no unfolding occurred, and emits a data-free fitted factor state that can score new observations on the same physical axes. - Docker images have explicit native, SCP, reference-acquisition, worker, and Hybrid control-plane roles with a common readiness endpoint.
- Hybrid clients, enrollment, Advisor relay, remote logging and protocol contracts have moved out of the OSS distribution. The public Workbench exposes neutral extension hooks, with no private-package discovery or activation from environment settings. Hybrid deployment procedures belong to the separately distributed private product.
The final checked candidate publishes the exact generated registry, managed profile, template, and scientific-authority counts. Counts in intermediate plans or screenshots are historical evidence, not a promise that every canonical node is eligible for managed optimization.
Breaking boundary cleanup
Version 0.6 removes prototype workflow/capsule readers, schema-1 evidence compatibility, arbitrary-estimator SDK validation helpers, retired SDK module exports, and deprecated dependency/configuration aliases. Current readers never silently translate those formats. See Migrate to 0.6 before upgrading a retained installation.
Canonical project packages now use the current schema-6 authority, and Campaign Review Packages use their own closed outer schema. Earlier project, campaign-evidence, and package schemas are refused instead of being silently upgraded. Re-export current work from the producing version before moving to 0.6, or recreate the workflow under the current explicit-validation contract.
For local, hosted, and Hybrid product choices, see Choose Your SpectraSherpa Path.