Research Library
Discovering a Volatility-State Signal in Model Innovations
How a small live volatility-dashboard experiment evolved into historical validation of innovation RMS, including warmup distortions, failed lead-lag claims, composite-index experiments, and the move toward asset-specific telemetry.
Research Status
Historical case study with a retained production measurement.
Innovation magnitude remains part of the current Vyreon production system. This article documents how that measurement was discovered, misinterpreted, repeatedly retested, and eventually reduced to a narrower and more defensible claim.
For the current definition and public interpretation of the signal, see Innovation Magnitude as a Volatility-State Signal.
Why This Article Exists
The present production article explains what the innovation-based volatility signal is and how it is used today.
This article explains how the research arrived there.
That distinction matters because the original volatility research contained several competing interpretations:
- Innovation dispersion as unexpected volatility.
- Innovation dispersion as expectation instability.
- Innovation slope as a leading volatility signal.
- Cross-horizon dispersion as a term structure of market stress.
- A composite Market Stress Index.
- A possible relationship with VIX, VVIX, SKEW, realized volatility, backwardation, and drawdowns.
- A possible argument for pooled multi-asset models.
- A later argument for asset-specific models.
Some of those ideas survived.
Some were useful only as exploratory framing.
Some did not survive historical testing.
The durable result was simpler:
The magnitude of causal model innovations is a useful volatility-state measurement, while changes in that magnitude can provide transition or hazard information without establishing a stable directional or fixed lead-lag signal.
The Starting Point
Vyreon was not originally designed as a volatility model.
The production system estimates expected forward returns across four maturity buckets:
- 8 to 30 calendar days.
- 31 to 60 calendar days.
- 61 to 120 calendar days.
- 121 to 365 calendar days.
When a horizon outcome matures, the system compares the realized value with the expectation formed before that outcome was known.
innovation
= realized horizon outcome
- prior expected horizon outcome
The innovation is a forecast error, but the word error can be misleading.
A large innovation does not automatically mean that the model is defective. It can also mean that the market departed sharply from the prior estimated state.
The original research question became:
If the expectation model is meaningful, does the magnitude of its innovations describe something real about market volatility and instability?
The First Dashboard Idea
The earliest volatility dashboard concept was intentionally conservative.
The proposed architecture kept the four horizon measurements separate, then added a secondary volatility context layer and a slow macro context layer.
The intended hierarchy was:
- Expectation stability, measured from model innovations and cross-horizon behavior.
- Volatility context, including compression, expansion, persistence, and term structure.
- Macro context, used only as background annotation.
This produced an important design rule:
Macro data may contextualize a measured market state, but it should not be used to invent an explanation for that state.
Federal Reserve Economic Data was considered for slow background variables. It was never required for the core telemetry measurement.
The dashboard was supposed to answer a few narrow questions:
- Which horizon became less stable?
- Which horizon remained coherent?
- Was instability local or spreading?
- Was volatility compressing, expanding, or stalled?
- Did cross-horizon agreement improve or deteriorate?
This was useful product thinking, but it did not yet establish what the innovation signal measured.
Early Live Observations
The first live window contained only a small number of observations.
Several visual patterns appeared compelling:
- Innovation dispersion seemed to change around volatility events.
- The level of innovation dispersion sometimes moved differently from VIX.
- The first difference of dispersion sometimes aligned more closely with changes in VIX.
- Longer-horizon and shorter-horizon measurements did not always move together.
- A composite index appeared to highlight periods of market stress.
These observations produced an early working interpretation:
innovation level -> current instability state
innovation slope -> instability transition
innovation acceleration -> shock onset
term slope -> long-horizon versus short-horizon instability
The interpretation was coherent enough to justify historical testing.
It was not yet strong enough for a production claim.
Why The Early Numbers Were Unstable
The source record contains materially different correlation estimates across experiments.
That was not one result changing randomly. The experiments were measuring different objects under different conventions.
The major sources of variation were:
Small Live Samples
Some early tests used only several weeks of live observations.
With roughly 30 to 50 usable rows, one event can dominate a correlation or lead-lag result.
A high correlation in that setting is evidence to investigate, not a stable parameter estimate.
Warmup Contamination
The earliest part of a burn-in includes initialization effects.
At different stages of development, these included:
- Kalman coefficient-state convergence.
- Rolling baseline initialization.
- Exponential normalization warmup.
- Incomplete maturity history.
- An arbitrary identity initialization for coefficient covariance, later replaced with a trained covariance.
Including those rows can make the signal appear unusually high, unusually smooth, or artificially predictive.
Pooled Versus Single-Asset Analysis
A pooled model combines assets with different:
- Options liquidity.
- Strike density.
- Volatility skew.
- Dealer positioning.
- Event risk.
- Return distribution.
- Options participation.
A SPY-specific signal can align tightly with SPY realized volatility and VIX-family measurements because they describe the same underlying market.
A pooled signal describes a broader cross-asset state and should not be expected to match SPY-specific benchmarks as closely.
Smoothing Span
A short exponential average reacts quickly.
A longer average tracks slower realized-volatility windows more closely.
Changing the smoothing span changes both correlation and apparent timing.
Target Construction
Close-to-close volatility, Parkinson volatility, VIX, VVIX, backwardation, SKEW, and forward absolute return are not interchangeable targets.
They describe different parts of volatility behavior.
Trailing-Window Effects
A realized-volatility series computed over the prior 21 days already contains persistence and delay.
Shifting it forward can create an apparent lead even when both measurements describe the same regime contemporaneously.
This became one reason the later research rejected a stable fixed lead-lag claim.
Historical Reconstruction
The decisive step was to recreate the innovation measurement over the longer walk-forward history.
That allowed the research to examine the signal across multiple environments, including:
- Low-volatility periods.
- The 2018 volatility shock.
- 2019 instability.
- The 2020 COVID transition.
- Post-2020 compression.
- The 2022 volatility regime.
- The 2023 and 2024 periods.
The historical plots showed that the innovation measurement was not merely a short live-data artifact.
The signal repeatedly expanded during periods of elevated realized volatility and compressed during calmer periods.
The relationship was visible across both close-to-close and Parkinson volatility.
The exact correlation depended on the implementation and validation convention, but the regime correspondence survived.
This changed the research question.
The question was no longer:
Is there any volatility-related information in the innovations?
The question became:
What is the narrowest interpretation that survives changes in period, smoothing, target, and implementation?
What Survived
1. Innovation RMS Is A Volatility-State Measurement
The most durable result is the level relationship.
A smoothed root-mean-square aggregation of innovations tracks broad realized-volatility regimes strongly.
innovation RMS
= sqrt(mean(innovation_i^2))
This is not surprising after the fact.
A return process can be written conceptually as:
realized outcome
= expected component
+ innovation
When the expected component is comparatively smooth, changes in the innovation process contribute heavily to realized dispersion.
The important evidence is not that one mathematical identity exists.
The important evidence is that the measured innovation series remains:
- Causal.
- Structured through time.
- Stable enough to smooth.
- Strongly associated with independent realized-volatility estimates.
- Coherent across historical and live periods.
2. Raw Level And Transition Are Different Measurements
The level of innovation RMS and its derivative should not be merged conceptually.
The level answers:
How large have recent departures from prior expectations been?
The derivative answers:
Is that departure magnitude increasing or decreasing now?
This supports the working distinction:
RMS level -> volatility state
RMS change -> transition or hazard state
The derivative is noisier and less stable.
Its most defensible interpretation is not a guaranteed forecast. It marks periods in which the distribution of future volatility conditions may differ from baseline.
3. Stress-Event Rates Changed After Large Telemetry Shocks
Internal event studies found materially different conditional rates after large changes in the telemetry signal.
Historical event tests found materially higher rates of volatility stress following large telemetry shocks, although the exact lift depended on the event definition and sample construction.
This supports continued research rather than a fixed operational conclusion.
They do not establish an immutable future probability.
The event definition, overlapping observations, threshold choice, and dependence structure all matter.
4. The Signal Is Non-Directional
Innovation magnitude predicts neither upward nor downward price resolution by itself.
A positive surprise and a negative surprise both increase RMS.
This is exactly what a volatility-state measurement should do.
5. Realized Volatility Is The Primary Benchmark
VIX and VVIX are useful comparison series, but they are market prices of volatility risk.
Realized volatility describes what price actually did.
The strongest validation therefore compares innovation RMS with independent realized-volatility estimators.
This is why the public chart uses both close-to-close and Parkinson volatility.
What Did Not Survive
A Fixed Lead-Lag Claim
Several early analyses appeared to show that telemetry led VIX or realized volatility by a small number of days.
That claim did not remain stable under broader historical testing and alternative alignment conventions.
The current research does not claim a fixed lead or lag.
The Claim That The Signal Is Simply Unexpected Volatility
The phrase is useful intuition, but it is incomplete.
Innovation magnitude includes:
- Market movement.
- Prior expected state.
- Coefficient uncertainty.
- Observation variance.
- Model adaptation.
- Horizon construction.
The signal is best described as model-state mismatch or innovation magnitude.
Calling it unexpected volatility can help explain the idea, but it should not be treated as a complete mathematical identity.
The Composite Market Stress Index As The Primary Observable
Several experiments combined:
- RMS level.
- RMS slope.
- Acceleration.
- Shock mass.
- Cross-horizon term slope.
The resulting index often looked similar to normalized RMS because RMS dominated the composite.
The project now prefers the primitive measurements over an opaque score.
A reader can inspect:
- Raw RMS.
- Smoothed RMS.
- Change in smoothed RMS.
- Cross-horizon state.
That is easier to audit than one weighted index.
Macro Variables As Explanatory Inputs
FRED and macro events can annotate the history.
They should not be used to explain every movement in the telemetry signal.
The system measures market state from its own inputs.
A later human narrative is not part of the measurement.
The Claim That High Correlation Proves Return Alpha
A useful volatility measurement does not prove that the expected-return signal produces a profitable trading strategy.
Calibration, state tracking, volatility correspondence, directional accuracy, execution, and profitability are separate questions.
The Per-Asset Model Lesson
The volatility research also changed the proposed scaling architecture.
A pooled model was valuable for testing whether the feature framework generalized across assets.
However, asset-level telemetry quality depends on the quality of the asset-level expectation model.
Different assets have different:
- Liquidity structures.
- Options participation.
- Strike grids.
- Skew behavior.
- Gamma and vega distributions.
- Event sensitivity.
This suggests a future architecture with:
- A shared causal feature pipeline.
- Asset-specific Ridge priors.
- Asset-specific Kalman states.
- Per-asset telemetry payloads.
- Cross-asset aggregation performed downstream.
In that design, systemic stress is constructed from clean asset observers rather than by forcing one pooled model to represent every asset identically.
This remains a future architecture, not the current public SPY implementation.
The Current Product Interpretation
The current public report uses innovation magnitude as a direct measurement.
It does not publish a Market Stress Index.
It does not claim a fixed lead over VIX or realized volatility.
It does not use FRED data to explain the state.
The report shows:
- Raw innovation dispersion.
- Smoothed innovation dispersion.
- Whether the raw value is above or below the smoothed trend.
- Whether the smoothed trend is rising or falling.
- Historical comparison with realized volatility.
This allows descriptions such as:
- Expanding instability.
- Compressing instability.
- Current relief inside a still-rising background.
- Renewed disturbance inside broader compression.
These are descriptions of measured relationships.
They are not promises about what happens next.
Why The Research Process Matters
This research is a useful example of why a discovery should not be frozen at its first compelling chart.
The initial live dashboard produced several plausible stories.
Historical testing then separated them:
- Strong level correspondence survived.
- A stable fixed lead did not.
- Event-rate changes remained promising.
- Composite scoring added little transparency.
- Per-asset calibration became more compelling than pooled inference for high-fidelity telemetry.
- Realized volatility remained the most important external benchmark.
The surviving measurement became narrower, but more trustworthy.
Current Conclusion
The volatility research produced one of Vyreon's strongest retained measurements.
The defensible conclusion is:
Causal innovation magnitude from the multi-horizon expectation model behaves as a volatility-state measurement. Its change contains useful transition information, but it does not provide a stable directional signal or fixed historical lead over realized volatility.
That conclusion is less dramatic than several early interpretations.
It is also better supported.