Supported File Types
File support is determined by the installed native ingestion registry. Optional algorithm extras do not change which source files can be parsed.
Base Install
These formats are handled without SpectroChemPy:
| Extension | Typical Use | Notes |
|---|---|---|
.csv |
tabular spectra, feature matrices, or one unheaded X/Y spectrum | Headered and unheaded layouts support comma, semicolon, or tab delimiters plus decimal-point or decimal-comma numbers. Ambiguous unheaded X/Y files require one visible, persisted interpretation. Keep sample IDs and target columns explicit for tables. |
.jdx, .dx |
JCAMP-DX spectra | Open spectroscopy interchange format. |
.npy, .npz |
NumPy arrays or SpectraSherpa synthetic datasets | .npz is also used for SpectraSherpa synthetic benchmark datasets. |
.mat |
MATLAB arrays and Eigenvector PLS_Toolbox DataSet Objects | MATLAB v4/v5 and v7.3/HDF5 numeric workspaces are supported. The qualified DSO contract preserves n-D data, per-mode scales, labels, titles, class sets, include masks, image layout, history, and descriptive metadata across both storage families. A two-dimensional pixels-by-features image DSO exposes both its ordinary unfolded table and an explicitly named, column-major image-cube view; the source inclusion mask is retained on both. Arbitrary MATLAB classes, executable objects, sparse/complex arrays, and ragged batch DSOs refuse. |
.opus, numeric suffixes such as .0 |
Bruker OPUS one-dimensional measurements | Every typed result is inventoried; select the exact scientific result for multi-result files. |
.spc |
Galactic/Thermo GRAMS SPC spectra and multifiles | Common-axis multifiles become one sample matrix; independent XYXY records remain explicit assets. |
.txt with #Wave / #Intensity tab header |
Renishaw WiRE single-spectrum text export | The exact two-column export is supported; other .txt tables, series, and maps are not guessed. Prefer WDF when acquisition metadata or topology is required. |
Native Bruker OPUS
One-dimensional Bruker OPUS (.opus and numeric suffixes such as .0) is
read natively. A file can contain absorbance, transmittance, reflectance,
interferogram, phase, or other typed results; the Workbench inventories them
and saves the scientist's exact result choice in the canonical DAG. Series and
3-D OPUS blocks are recognized and refused rather than flattened.
The always-running bundled corpus contains one licensed real acquisition. Additional qualified multi-asset layouts are exact-hash external references because their source corpus does not grant redistribution permission; those bytes must be acquired and executed during release qualification. Support is limited to the explicitly qualified grammar and is not inferred from the numeric filename alone. The optional SpectroChemPy differential job compares the retained OMNIC and WDF readers; it is not an OPUS or SPC variant oracle.
Missing spectral values
Ingestion preserves an empty numeric cell or an explicit supplier missing token
such as #NaN as missing data so source identity is not silently rewritten.
Binary missing values are narrower: the qualified OMNIC SPA reader preserves
them only for the retained code-23 profile when the missing-value mask exactly
matches one explicit processing-history Blank ... From ... to ... interval.
An arbitrary SPA NaN, a mismatched blank interval, or any infinite coordinate
or intensity is refused.
The admitted file remains inspectable and carries a visible missing-data
warning, but numerical preprocessing, variable selection, transfer, fitting,
and prediction remain fail-closed: SpectraSherpa does not invent replacement
intensities. For a terminal vendor-declared blank interval, apply the canonical
preprocess.clip_range operation to retain only the measured region before any
numerical preprocessing. Otherwise exclude identified affected samples or
variables, or repair the source. Automatic or fitted imputation is not a 0.6
capability; it requires a separately reviewed preprocessing contract with
training-only fit semantics and explicit provenance.
Native Galactic SPC
Galactic/Thermo GRAMS SPC (.spc) is read through the native bounded parser.
Qualified 0x4B little-endian files cover generated, common explicit, and
independent XYXY axes; qualified 0x4D files cover the old word-swapped data
layout. Common-axis multifiles become one chemometric sample matrix. XYXY
files retain every independent record as a named asset and require an exact
selection. Big-endian 0x4C remains fail-closed until an independent scientific
conformance fixture is available.
Native Thermo OMNIC and Renishaw WiRE
Qualified legacy Thermo OMNIC SPA single spectra, SPG compatible groups, and
SRS time series (.spa, .spg, .srs) are read by the native registry. SRS
coverage includes the qualified rapid-scan, high-speed, GC/TGA, and extended-
directory Kinetics layouts; it is not a claim that every historical SRS layout
is interchangeable. SPA projection retains the declared ordinate type, axis,
acquisition timestamp and scan fields, plus bounded comments, processing
history, experiment information, accessory text, and the identity—not the raw
values—of a qualified embedded linked source when those records are present.
For the retained Nicolet Apex code-23 profile, OMNIC displays Transmittance
while processing history separately records Final format: Single Beam.
SpectraSherpa preserves both statements, assigns no ratio units, and therefore
does not offer transmittance-to-absorbance conversion for that profile: a
display label alone does not establish an I/I0 ratio. The declared blanked
interval is preserved as missing in SPA and CSV, omitted by the corresponding
JDX export, and filled by the corresponding SPC export. Cross-format signal
parity for those sources is claimed only over their common measured finite
region.
Metadata-poor CSV, JCAMP-DX, or SPC exports are not used to invent fields that
the export discarded.
Qualified Renishaw WiRE WDF (.wdf) single spectra, depth series, XY lines,
StreamLine acquisitions, and rectangular maps are also native. Map spectra
remain one sample-by-Raman-shift matrix with exact X/Y/Z/time origin columns
and explicit topology. The exact tab-separated WiRE single-spectrum text
export headed by #Wave and #Intensity is also native; it retains only the
Raman-shift axis and intensity, so its result warns scientists to use WDF when
acquisition settings, coordinates, or map topology are required. Thermo
Paradigm/OMNICxi containers (.srsx,
.session, .map, .mapx) still require an upstream spectrum export.
The Workbench lists recognized pending families as unavailable and shows the
same export remediation before a file is stored. It never sends them through
a generic parser or silently guesses other .txt/.dat content.
All admitted source formats—CSV, JCAMP-DX, NumPy, MATLAB v4/v5 and v7.3, the qualified Eigenvector PLS_Toolbox DataSet Object contract, Galactic SPC, one-dimensional Bruker OPUS, qualified legacy OMNIC SPA/SPG/SRS, and Renishaw WiRE WDF and exact single-spectrum text exports—are parsed by bounded native readers with no SpectroChemPy import or fallback, proven per PR by a clean-room no-SpectroChemPy profile. SpectroChemPy remains an optional runtime for exactly three scientific operations: EFA, MCR-ALS, and SIMPLISMA.
This is SpectroChemPy-independent ingestion complete for the qualified format
set, not complete spectroscopic ingestion. Additional vendor and interchange
families remain demand-gated, and Thermo Paradigm/OMNICxi containers (.srsx,
.session, .map, .mapx) remain pending with the export guidance above.
Optional SpectroChemPy Algorithms
0.6.0 release lifecycle. Install 0.6.0 from PyPI only after the public index reports that exact version. Before the public tag exists, use only the exact monorepo commit named by the qualification record. After the tag exists but before PyPI reports 0.6.0, use the exact
spectra-sherpa-v0.6.0source tag. Source version text alone is not publication evidence.
After 0.6 publication, install with:
pip install "spectra-sherpa[scp]==0.6.0"
This enables the three SpectroChemPy-backed
operations retained by SpectraSherpa—EFA, MCR-ALS, and SIMPLISMA. It does not
install a SpectraSherpa reference-data catalog. SpectraSherpa itself currently
supports Python >=3.11,<3.13; the optional dependency is pinned to exactly
SpectroChemPy 0.8.1. Installing this
extra does not add source-file readers.
Optional HITRAN/HAPI Extra
Install with:
pip install "spectra-sherpa[hitran]==0.6.0"
This enables HITRAN/HAPI synthesis support. The current package supports hitran-api >=1.3.0.0,<2 and hitran-api2 >=0.2.2,<1. See the official HITRAN HAPI page and HAPI manual for upstream API details.
Practical Guidance
For a first demo, prefer CSV, JCAMP-DX, NPY/NPZ, a qualified MATLAB/DSO
workspace, qualified SPC,
qualified one-dimensional OPUS, qualified legacy OMNIC, qualified WDF, or the
exact Renishaw #Wave/#Intensity text export. For a pending native instrument
format, export a representative source to an admitted format and verify its
dimensions, axis, and units before converting a large calibration library.
After import, check the Files, Metadata, and Data Matrix panels to confirm the file names, extensions, sample count, spectral axis, and target values. For an n-dimensional DSO, the Data Matrix panel instead shows the full native shape and dimension roles. Inspect every axis, class, and include set. Use Dimension Projection before a two-dimensional node. For a qualified image DSO, select its explicitly named image-cube result before manually adding PARAFAC; masked PARAFAC honors the source's spatial exclusions and records that decision. The Workbench never silently flattens or refolds a DSO.