The Sound of Risk: Acoustic Signals for Volatility Forecasting

Are you struggling to find alpha in crowded factors? It might be time to expand your creative horizon.
When analyzing earnings calls, the first instinct for many is to examine the text transcripts.
What did the company say? Which words did management use? Was the tone positive or conservative? Can Large Language Models (LLMs) extract sentiment signals from them? This direction is certainly valuable, but it faces a realistic problem: text can be carefully curated.
In particular, public communications from listed companies are often polished in advance—smooth, steady, and watertight.
Even if a company is facing significant operational difficulties, senior management usually feels compelled to convey a positive and optimistic message. Sometimes this is unavoidable because if the company admits its struggles or undermines itself, it can trigger a chain reaction with unacceptable consequences.
Thus, we often see a company, despite planning a major technological transformation, stubbornly insisting that "our existing technology is断层领先 (断层领先 means 'leading by a wide margin' or 'unmatched')."
So, how do we capture these overtones?
Today, I introduce a new perspective from the paper: The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability.
The authors shifted the paradigm: since "what is said" can be packaged, why not analyze "how it is said"?
More specifically, instead of just analyzing the text content of earnings calls, they analyzed the vocal signals of executives, such as tension, fluctuation, stability, and excitement—clues that are harder to fake.
This is the most compelling aspect of reading this paper. It asks a profound question:
If the script itself increasingly resembles a PR product, what should the market truly listen to?
The answer provided by the paper is not in the text, but in the voice.
1. Why Research This Direction?
The first reason is that traditional text analysis has been nearly exhausted.
Earnings calls, shareholder letters, and press releases have been swept by Natural Language Processing (NLP) in recent years. From sentiment dictionaries to BERT, and now to LLMs, everyone is studying "what management said." This direction is certainly valuable, but it has an insurmountable problem:
Public text is the easiest to refine.
The prepared remarks at the beginning of an earnings call are products repeatedly polished by legal, Investor Relations (IR), and management. You can read attitudes from them, but you cannot guarantee you are reading the "true state."
The second reason is that earnings calls have a natural "moment of truth."
That moment is the Q&A.
Anyone can stay steady when reading from a script. When analysts ask follow-up questions, stability is not guaranteed. Especially when asked about guidance, inventory, demand, gross margins, regulation, or capital expenditures—points prone to risks—executives' reactions often shift from "prepared expressions" to "improvisation."
At this point, the text may maintain decorum, but the voice may not cooperate.
The third reason is that the technical conditions have matured.
If this were ten years ago, researchers would likely only rely on manual acoustic features like MFCCs, pitch, energy, and pauses, which would be useless in high-noise environments. However, with the rise of next-generation speech models like wav2vec 2.0, Conformer, and WavLM, researchers now have the opportunity to stably learn higher-level representations from low-quality phone audio.
In other words, people didn't previously doubt that "there might be something in the voice"; rather, the intuition existed, but the tools were insufficient. Now, the tools have finally caught up.
So, what is this paper most like?
It is not that the authors suddenly had a flash of inspiration and invented a mystical direction; rather, several conditions coincided at the same time:
- Text signals are becoming increasingly crowded.
- The Q&A scenario naturally carries pressure-testing attributes.
- Speech foundation models can finally handle such tasks.
The combination of these three events transformed "finding risk in voice" from a casual intuition into a researchable problem.
This is also an important source for quant researchers to maintain their career longevity: knowing which directions are valuable, quietly waiting for the paradigm shift brought by technological maturity, and seizing the opportunity first.
2. The PIAM Acoustic Model
The paper studied 1,795 earnings calls from 283 Nasdaq companies between 2018 and 2023, totaling nearly 1,800 hours. It combined three types of data:
- Raw audio
- Transcribed text
- Financial market data
Then, it built a multimodal framework, core named PIAM (Physics-Informed Acoustic Model), or "an acoustic model with physical constraints."
PIAM's underlying architecture is not just a standard speech recognition model; it introduces the Westervelt equation from nonlinear acoustics. The essence of this approach is: when management is under pressure or attempting to conceal information, the vocal cords exhibit nonlinear physiological jitter, which superimposes on normal speech signals.
By using physical laws as a regularization term, PIAM acts like a 'microscope' to remove artifacts caused by telephone line compression and microphone clipping, precisely capturing the subtle physiological features of 'insincerity.'

To explain this method in non-academic terms, it takes just three steps to "put an elephant in a refrigerator."
Step 1: Extract representations from raw speech.
This approach is similar to common self-supervised speech models today: first, let the model "listen" to enough sounds, then learn to map sounds into informative vector representations. The paper uses a思路 (approach) similar to wav2vec 2.0, combined with Bi-LSTM and attention mechanisms, to capture the most critical segments of a sentence.
Step 2: Not just transcription, but simultaneous judgment.
PIAM outputs three types of results simultaneously:
- Transcribed text
- Vocal emotion labels
- Acoustic event labels
Things like silence, laughter, coughing, and changes in call quality are also identified as much as possible. Because in earnings calls, these are not "dirty data"; often, they are clues in themselves.
Step 3: Map vocal emotions and text emotions into the same coordinate system.
The paper does not stop at discrete labels like "happy," "nervous," "angry," or "afraid." Instead, it maps both sides into a unified three-dimensional emotion space: ASL (Affective State Label), with three dimensions:
Tension: Tension levelStability: Stability levelArousal: Activation level
The practical benefit of this approach is clear. Financial modeling dislikes vague statements like "this sentence feels like fear." It prefers:
How much did the CFO's stability mean decrease during the Q&A phase?
Once emotions are mapped to continuous variables, you can continue to calculate means, volatility, skewness, kurtosis, and, most importantly, phase-switching deltas.
The paper also provides an interesting distribution chart showing the difference in emotional expression between voice and text for different roles:

The CEO is always the person setting the highest tone in the company, always more positive. For example, taking anger as an instance, although the text analysis shows an anger value of only 0.1, the voice analysis reveals an actual anger value of 0.8. Taking happiness as another example, while the spoken words might seem happy, the voice analysis shows they are not that optimistic.
Interestingly, the CFOs show relatively consistent emotional values whether analyzed through text or voice. This is precisely where the model's value lies: once the CFO's vocal emotion diverges from the text emotion, information entropy surges—only low-probability events have dissemination value.
This is also reasonable. CEOs often handle visions, directions, and narratives; CFOs are more likely to be challenged in Q&A with questions that cannot be handled by narrative alone, such as inventory, cash flow, profit margins, accounting treatments, capital expenditures, and guidance fulfillment.
3. Capturing the Cracks in Voice
The smartest cut in this paper lies in the Q&A section.
The authors did not treat the entire earnings call as a unified text block to calculate emotion. Instead, they carefully distinguished two segments:
- Prepared remarks
- Analyst Q&A
This segmentation is crucial.
Because these two segments are fundamentally different types of information.
The former is more like a press conference, while the latter is more like a stress test.
The former looks at "what the company wants to convey," while the latter looks at "whether the company remains steady when pressed."
Thus, the paper focuses on an interesting aspect:
Does the executive's emotion change significantly when switching from scripted remarks to impromptu Q&A?
This is far more interesting than "whether the average emotion of the entire earnings call is positive or negative."
The average value is too easily smoothed out by the first half. What is truly informative is often the fluctuation at the moment of the switch.
4. Quantified Emotional Volatility
So, is this article a prediction model? It's not that simple.
Its most important conclusion is not "executives are nervous in voice, so stock prices will fall," but a finer, more realistic conclusion:
Emotional signals in voice and text have little help in predicting "future return direction," but significant help in predicting "future volatility."
This point is critical.
Many people, upon seeing such research, intuitively think: "Can we predict tomorrow's price direction by listening to earnings calls?"
The paper's answer is basically negative.
But it excels in another area: predicting whether the market will become more uneasy in the future.
Specifically:
- For future return direction, the model has almost no stable predictive power.
- For future volatility, the model performs significantly better.
- When predicting 30-day realized volatility, the out-of-sample $R^2$ of the complete multimodal model reaches
0.438. - The traditional financial factor baseline model is approximately
0.251.
This gap is not small.
If we translate this into plain language:
Listening to how executives speak may not tell you "whether the stock will rise or fall tomorrow"; but it likely tells you "whether this stock will be more volatile in the near future."
This is why I feel this paper is not mystical.
Direction is inherently difficult to predict due to too many exogenous variables; but "whether uncertainty is rising" is often easier to leak from management's state.
The market is often not afraid of bad news, but afraid of not knowing how much bad news hasn't been said yet.
From this perspective, the paper captures not "bad outcomes," but the premature heating up of uncertainty itself.
For options traders, being able to predict changes in volatility is exactly the "holy grail" they seek. Price direction is a "vector," while volatility is a "scalar." In financial derivatives markets, scalars are directly tradable assets. The value of the PIAM model lies in providing a "measure of uncertainty" that is earlier and more accurate than market consensus.
5. Physical Constraints
The second step of the model, physical constraint regularization, is the most technical part of the paper and the part most likely to lose readers, but the underlying intuition is not complex.
The authors argue that the problem with earnings call audio is not just ordinary noise, but more commonly nonlinear distortion, such as:
- Microphone overload causing sound clipping
- Aggressive telephone system compression deforming details
- Low-bitrate transmission introducing artifacts
If you treat all of these as "random noise," the model might learn the quirks of the telephone system as emotional features.
Therefore, the paper borrows the Westervelt equation from nonlinear acoustics and uses it as a regularization term to constrain the model's latent representations from becoming too absurd.
In one sentence:
Voice is not pure digital data; it involves the physical processes of sound production and propagation. Since the distortion in earnings call audio follows patterns, the model should ideally possess some physical common sense.
This idea is not new in fluid dynamics, meteorology, or materials science, but it appears novel in the context of earnings calls.
6. Silent Burst
What truly fascinates me about this paper is not the multimodality or the term "physics-informed" itself.
It is that it captures the most awkward yet realistic aspect of the earnings call scenario:
**This is a highly performative occasion, but it is impossible