Configuration
Copy examples/kinextract.config to your extraction directory and edit
the galaxy-specific parameters:
[wavelength]
zgal = 0.00157811 # galaxy redshift
wavefitmin = 8400.0 # fit range (Angstrom)
wavefitmax = 8800.0
[kinematics]
xlam_auto = true # auto-select regularization
xlam_smooth_threshold = 0.25
[continuum]
fit_continuum = true
TOML section names are purely organizational – every field is flattened
into a single FitConfig namespace regardless of
which section it’s under. examples/kinextract.config documents the most
commonly-tuned fields inline; for the complete list, call
FitConfig.describe() or read the class docstring below.
fit_continuum = true turns on continuum co-fitting: a penalized-B-spline
continuum is folded directly into the same optimization as the LOSVD and
template weights (see kinextract.joint). Alternatively, pre-normalize
the spectrum once (e.g. via the standalone
kinextract.continuum.asymmetric_least_squares_continuum() utility)
and fit with fit_continuum = false.
- class kinextract.config.FitConfig(gal_file: str = '', template_list_file: str = 'Tlist', template_dir: str | None = None, template_npz_file: str | None = None, template_npz_select: list[tuple[float, float]] | None = None, regions_bad_path: str = 'regions.bad', outdir: str | None = '.', write_outputs: bool = False, wavemin_full: float | None = None, step: float | None = None, wavefitmin: float = -inf, wavefitmax: float = inf, zgal: float = 0.0, norm_wave_frame: str = 'auto', galaxy_params_path: str | None = None, use_galaxy_params_velocity_bounds: bool = True, sigl: float = 100.0, xlam: float = 300.0, xlam_auto: bool = False, xlam_auto_grid: tuple = (100.0, 1000.0, 10000.0, 100000.0, 1000000.0, 10000000.0), xlam_criterion: str = 'discrepancy', xlam_chi2_tolerance: float = 0.02, xlam_smooth_threshold: float = 0.25, xlam_discrepancy_nsigma: float = 0.3, xlam_max_peaks: int = 1, xlam_peak_min_prominence: float = 0.1, xlam_auto_maxiter: int | None = None, xlam_wing_shrink: float = 0.0, xlam_wing_shrink_sfac: float = 1.8, losvd_vmin: float | None = None, losvd_vmax: float | None = None, n_losvd_bins: int = 29, fit_continuum: bool = False, joint_n_interior_knots: int = 10, joint_degree: int = 3, joint_xlam_cont: float = 3.0, joint_cont_diff_order: int = 2, joint_n_sigl0_iter: int = 3, joint_sigl0_tol: float = 2.0, joint_recenter_v: bool = True, joint_prenorm: bool = False, norm_error_mode: str = 'unit', fortran_rebin_norm: bool = True, norm_apply_fortran_flux_mask: bool = True, norm_fortran_flux_min: float = 0.0, norm_fortran_flux_max: float = 1.3, spec_col3_is_variance: bool = False, use_spectrum_errors: bool = True, mask_emission_line_velocity_pad_kms: float = 300.0, extra_absorption_lines: tuple = (), icoff: int = 1, coff: float = 0.0, coff2: float = 0.0, fortran_nlosvd_full_x: bool = True, fortran_template_mixture: bool = True, fortran_mask_template_outside: bool = True, fit_global_amp: bool = False, continuum_poly_mode: str = 'none', continuum_poly_bound: float = 0.1, template_w_bounds: tuple | None = None, data_fwhm_A: float | None = None, template_fwhm_A: float | None = None, data_fwhm_frame: str = 'observed', map_maxiter: int = 10000, map_maxfun: int = 200000, map_ftol: float = 1e-08, map_gtol: float = 1e-08, map_maxls: int = 50, use_scaled_optimizer: bool = True, use_jax_objective: bool = True, jax_enable_x64: bool = True, print_every: int = 200, nuts_num_warmup: int = 50, nuts_num_samples: int = 75, nuts_num_chains: int = 4, nuts_seed: int = 0, clean: bool = False, clean_sigma: float = 3.0, clean_maxiter: int = 5, clean_minpix: int = 10, clean_bloom_pixels: int = 0, clean_protect_ca_triplet: bool = True, clean_ca_centers: tuple = (8498.02, 8542.09, 8662.14), clean_ca_half_width: float = 6.0, clean_protect_velocity_pad_kms: float = 400.0, clean_protect_absorption_only: bool = False, clean_protect_ca_frame: str = 'rest', clean_protect_windows: tuple = (), clean_protect_windows_frame: str = 'rest', mask_emission_lines_in_fit: bool = True, mask_emission_line_half_width_A: float = 5.0, mask_emission_line_snr_threshold: float = 3.0, mask_emission_line_snr_context_A: float = 20.0, mask_emission_grow_A: float = 0.0, segment_emission_mask: bool = True, segment_emission_n_sigma: float = 3.0, segment_emission_win_A: float = 50.0, mask_paschen_lines_in_fit: bool = False)[source]
All tunable parameters controlling a single LOSVD spectral fit. Run FitConfig.describe() to see a grouped, filterable table of all fields and their one-line descriptions.
FitConfigis the single configuration object passed torun_spectral_fit()andLOSVDErrorEstimator. Fields are grouped below by subsystem; calldescribe()at any time (on the class or an instance) to print this same grouping resolved against actual values, optionally filtered to a subsystem, e.g.cfg.describe("joint").Field groups
- Paths
Spectrum/template/output file locations.
- Wavelength/redshift
How raw pixel indices or wavelength columns map to rest-frame A.
- Kinematic grid
The non-parametric LOSVD velocity grid (
losvd_vmin/vmax,n_losvd_bins) and the fixed regularization strengthxlam.- Auto xlam selection
Grid-search selection of
xlaminstead of a fixed value.- Continuum mode
Pre-normalised input vs. co-fit P-spline continuum baseline (see
kinextract.joint).- Joint continuum, Pre-normalised mode
Options specific to each continuum-handling path.
- LOSVD/template
Template-mixture and continuum-offset conventions inherited from the legacy Fortran objective function.
- Optimizer
L-BFGS-B iteration/tolerance/JAX settings for the reported fit itself, and internally for the
xlamsearch andclean-mode outlier rejection.- Bayesian sampling
NUTS/HMC posterior-sampling settings, used only by the optional, non-default full-posterior path (
kinextract.bayesian.fit_state_bayesian(), called directly rather than throughrun_spectral_fit()).- Kinematic cleaning
Iterative sigma-clipping of outliers from the LOSVD chi-squared fit itself.
- Emission masking
Pre-fit exclusion of known and unknown emission-line pixels.
Performance
The reported fit is a bound-constrained L-BFGS-B optimum (MAP point estimate) of
chi2 + wing-tapered smoothness penalty + LOSVD normalization penalty, the same objective minimised by the original Fortran implementation this package is a port of – fast (typically a few seconds for a single-template mock or moderate real spectrum, including thexlamauto-selection grid search). A full-posterior (NUTS/HMC) alternative is available viakinextract.bayesian.fit_state_bayesian()for users who want a sampled posterior instead of a point estimate (see that function’s module docstring); it is not the default because a comprehensive recovery-accuracy comparison (LOSVD-shape L1 distance to a known truth, across MUSE/STIS at and away from their resolution limits, for both Gaussian and moderately non-Gaussian LOSVDs) found the MAP point estimate matches or exceeds the posterior mean’s shape recovery in every tested condition, at a small fraction of the runtime (seconds vs. tens of seconds to minutes for NUTS). Always checkresult["result"].success(the L-BFGS-B optimizer’s own convergence flag) – ifFalse, inspectresult["result"].messageand consider raisingmap_maxiter/looseningmap_ftolrather than trusting the point estimate.Known limitations: recovery near the instrumental resolution limit
Validated with a synthetic multi-instrument sweep (
benchmarks/; MUSE and STIS setups, both at and away from each instrument’s own resolution limit, for Gaussian and moderately non-Gaussian Gauss-Hermite LOSVD truths, 6 seeds per condition): once a real instrumental line-spread function is modeled and matched viadata_fwhm_A/template_fwhm_A(rather than assumed already-matched), recovery degrades gracefully near the resolution limit rather than failing outright, but a real, condition-dependent bias remains – there is no universal correction, which is exactly whyassess_recovery_bias()exists (see below).V-bias is real but not one-signed or uniform across instruments/resolution regimes/LOSVD shapes: in a 216-fit validation sweep (2 instruments x 3 resolution regimes x 3 LOSVD shapes x 2 S/N levels x 6 seeds), mean bias_V per condition ranged from about -5.2 km/s (MUSE, 1.5x LSF sigma, Gaussian truth) to about +8.2 km/s (MUSE, 4x LSF sigma – i.e. well-resolved – strongly non-Gaussian truth), with scatter of a similar magnitude. There is no single “expect X km/s bias” number that holds across setups, and bias does not simply vanish away from the resolution limit for strongly asymmetric LOSVDs (see the xlam_auto note below).
Sigma is more consistently overestimated, sometimes substantially: mean bias_sigma in the same sweep ranged from about -5.2 km/s (MUSE, well-resolved, Gaussian truth) up to about +13.3 km/s (STIS, 1.5x LSF sigma, strongly non-Gaussian truth) – consistent with the well-known effect of noise amplification when deconvolving a velocity dispersion at or below an instrument’s own LSF, not a bug or a tunable knob.
The MAP objective’s wing-tapered smoothness penalty is always zero-centered (matching the original Fortran convention), rather than recentered on a data-driven velocity estimate for each fit. Recentering can amplify bias whenever the velocity estimate itself is imprecise – a few to several km/s off, which is routine for a coarse cross-correlation – so a fixed zero point is the more robust default across a broad range of targets. Use
assess_recovery_bias()(below) to check whether a specific target near its instrument’s resolution limit would benefit from a more tailored treatment.bias_corrected_losvd()(analytic hat-matrix linearization) is a separate, older correction attempt that amplifies noise catastrophically (h3/h4 pinned at their fit bounds) when the LOSVD is this weakly identified. Do not use it as a resolution-limit correction.This is distinct from – and should not be confused with – trying to recover a velocity dispersion below the detector’s native pixel sampling with no LSF modeled at all (e.g. sub-pixel broadening): that regime showed much larger, less predictable bias in testing and is closer to an information-theoretic limit than a resolution-limit one.
Recommended practical guidance: since the direction and size of the bias is condition-dependent, don’t rely on the generic numbers above for a specific target – use
assess_recovery_bias()to measure the empirical bias directly, on mock spectra matched to that target’s own instrument resolution, template mixture, continuum, and noise level (built automatically from a completed fit viabuild_matched_mock()), then optionally applycorrect_recovered_losvd()to correct for it. This is only worth the extra runtime (one MAP refit per mock seed) when the target’s sigma is within about 2x the instrument’s LSF sigma; well away from the resolution limit the bias shrinks toward zero.xlam_auto(the default regularization-strength selector) can occasionally over-smooth strongly asymmetric (largeh3/h4) LOSVDs for specific noise realizations, biasing V by 15-30 km/s in roughly a quarter of tested cases for a strongly non-Gaussian truth – tighteningxlam_chi2_tolerancereduces this but was found to regress the common near-Gaussian case, so the default is unchanged. For targets suspected to have strongly asymmetric LOSVDs, cross-check withassess_recovery_bias()rather than assuming the auto-selectedxlamis well-calibrated. Insufficient regularization more generally (a fixed, lowxlamchosen by hand rather thanxlam_auto) reproduces the same failure mode and is not specific to strongly asymmetric truths – it is the standard bias-variance tradeoff of any penalized non-parametric fit, present in the original Fortran implementation too for an equally under-regularized choice, not a new behavior.Fit cost scales strongly with template count, independent of pixel count – e.g. 5 templates cost roughly 4-7x a single template, and a 35-template library roughly 12x (at fixed
xlam), mostly fromxlam_auto’s internal grid search needing a fresh JAX kernel compile per candidate value. This compounds multiplicatively withresidual_bootstrap()’s per-replicate cost. For large template libraries, consider a smaller, pre-selected subset, and setxlam_auto=False(with a pre-chosenxlam) for repeated refits (bootstrap andassess_recovery_bias()already do this internally where appropriate).
Known limitations: joint-mode (
cfg.fit_continuum=True) velocity biasValidated with a synthetic E-MILES sweep (2-SSP and broad, ~20-SSP mixtures;
n_losvd_bins=89– not the default, see that field’s own docstring – with a 1000A fit window; sigma=30-350 km/s, 3-5 seeds per condition):Sigma recovery is good across the full tested range (roughly +-1-3 km/s bias for sigma <= 200 km/s, growing to a still-modest +-1-4 km/s even at sigma=250-350).
V recovery carries two distinct biases. A small, roughly shrinkage-toward-zero offset (a few km/s, growing with the true velocity’s distance from the LOSVD grid’s own center at v=0) is present even at low sigma and is not yet understood – ruled out as the cause: xlam_criterion choice, wing-taper v_center recentering (including recentering on the exact true velocity, not just a cross-correlation estimate), fitting-template count, and generating-population complexity (a broad, ~20-SSP mixture shows the same pattern as a simple 2-SSP one). Separately, a larger bias (growing to roughly -10 to -14 km/s by sigma=350) appears at broad sigma; this component also appears, at comparable magnitude, in a parametric (Gauss-Hermite) fit of the same mock, consistent with the well-documented V/sigma/h3 covariance at low per-resolution-element S/N (van der Marel & Franx 1993; Cappellari 2017) rather than being specific to this package’s non-parametric LOSVD parameterization.
assess_recovery_bias()andcorrect_recovered_losvd()currently refuse joint-mode fits entirely (seeassess_recovery_bias()’s own guard) – there is no automated bias-correction path for this mode yet. Until the small offset above is root-caused, treat recovered V at low-moderate sigma as accurate to a few km/s, and rely onresidual_bootstrap()’s error bars (not the point estimate) at sigma >~ 200 km/s.
Examples
>>> from kinextract import FitConfig >>> cfg = FitConfig(template_list_file="Tlist", zgal=0.0016, ... wavefitmin=8400.0, wavefitmax=8800.0, ... fit_continuum=True, clean=True) >>> cfg.describe("clean")
- describe() None
Print every tunable field, grouped by subsystem, with a one-line description.
- Parameters:
pattern (str, optional) – Case-insensitive substring filter matched against the field name, its section name, or its description, e.g. FitConfig.describe(‘joint’).
Notes
Called on the class (FitConfig.describe()), shows dataclass defaults. Called on an instance (cfg.describe()), shows that instance’s actual current values instead.