snr_predictor
Author: Bagus Tris Atmaja (with Claude Code) Affiliation: NAIST Date: 2026.09
SNR recognition model of the Lombard machine speech chain (Novitasari et al., IEEE/ACM TASLP 2022, Fig. 2c). Given a noisy speech waveform it predicts the SNR class of the environment and produces the SNR embedding Z_SNR fed back to the TTS. It is made of stacks of convolution + residual blocks followed by linear layers.
ResBlock1d
Bases: Module
Two Conv1d-BatchNorm layers with a residual connection.
Source code in speechain/module/standalone/snr_predictor.py
SNRPredictor
Bases: Module
Predict the SNR class of noisy speech and produce the SNR embedding.
Input waveforms are turned into log-mel spectrograms by a built-in frontend, normalized per
utterance (SNR is a relative measure, so the absolute level is discarded), and encoded by
len(conv_dims) stacks of [Conv1d + ReLU + ResBlock]. Frame-level embeddings are produced by a
linear layer and averaged over the valid frames to obtain the utterance-level embedding. The
frame-level outputs allow short-term feedback in dynamic noise environments.
Source code in speechain/module/standalone/snr_predictor.py
41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 | |
extract_feat(wav, wav_len)
Turn waveforms into per-utterance-normalized log-mel features.
Source code in speechain/module/standalone/snr_predictor.py
forward(wav, wav_len)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
wav
|
Tensor
|
(batch, wav_maxlen, 1) waveforms or (batch, feat_maxlen, feat_dim) features. |
required |
wav_len
|
Tensor
|
(batch,) |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, Tensor]
|
Dict with logits: (batch, class_num) utterance-level SNR class logits emb: (batch, emb_dim) utterance-level SNR embedding Z_SNR frame_logits: (batch, feat_maxlen, class_num) frame_emb: (batch, feat_maxlen, emb_dim) feat_len: (batch,) number of valid frames |
Source code in speechain/module/standalone/snr_predictor.py
module_init(snr_classes, frontend=None, conv_dims=None, conv_kernel=3, res_blocks=1, emb_dim=256, dropout=0.1)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
snr_classes
|
List[str]
|
List[str] The names of the SNR classes, e.g. ['clean', '0', '-10']. |
required |
frontend
|
Dict
|
Dict The configuration of the waveform frontend (type + conf). If not given, the input of forward() must be acoustic features (batch, feat_maxlen, feat_dim). |
None
|
conv_dims
|
List[int]
|
List[int] The channel number of each convolution stack. |
None
|
conv_kernel
|
int
|
int Kernel size of the convolution layers. |
3
|
res_blocks
|
int
|
int Number of residual blocks in each stack. |
1
|
emb_dim
|
int
|
int Dimension of the SNR embedding Z_SNR. |
256
|
dropout
|
float
|
float |
0.1
|