2022/03/15

Python: Short-Time Fourier Transform (STFT) and its inverse transform (ISTFT) using librosa

Below shows an example using STFT to transform speech into the frequency domain and ISTFT to transform the spectral data back to waveforms in the time domain. The Librosa package is used.

import librosa

import soundfile as sf

x, sr = librosa.load('1.wav', sr=None)


n_fft = 256

hop_length = 128

win_length = 256


X = librosa.stft(x, n_fft=n_fft, hop_length=hop_length, win_length=win_length)


y = librosa.istft(X, n_fft=n_fft, hop_length=hop_length, win_length=win_length)


sf.write('1_istft.wav', y, 16000)

Results for x and y:



References:

Matlab: Short-Time Fourier Transform (STFT) and its inverse transform (ISTFT) (StudyEECC)

librosa.stft

librosa.istft

2022/03/10

Matlab: Short-Time Fourier Transform (STFT) and its inverse transform (ISTFT)

For sound processing in realtime scenarios, it is not possible to wait for a complete sound file. In this case, short-time processing is important.

Here is an example using STFT to transform speech into the frequency domain and inverse transform the spectral data back to waveforms in the time domain:


[S,F,T] = stft(x, fs,'Window',win2,'OverlapLength',128,'FFTLength',256); 

[y,ti] = istft(S,fs,'Window',win2,'OverlapLength',128,'FFTLength',256);

References:

stft - Short-time Fourier transform (MathWorks)

istft - Inverse short-time Fourier transform (MathWorks)

2022/03/09

Matlab: Discrete Cosine Transform (DCT) for speech processing

The DCT and inverse DCT may be used to convert a speech signal into the transform domain using real values and back to the speech waveform.

Example with Matlab:

%load speech

load mtlb

x = mtlb;

X = dct(x);

y = idct(X);

Results:


The values for x and y look the same, but MATLAB consider them as different.

>> isequal(x,y)

ans =

  logical

   0

For real-time processing, short-time DCT is required.

References:

Discrete Cosine Transform (Wikipedia)

dct (MathWorks)

2021/12/20

Math for Deep Learning: arg min

arg min f(x) = the value x with the minimum value of f(x)

這個arg min f(x)的意思是,當f(x)是最小值的時候,x的數值


e.g. 以下的例子

f(0) = 3

f(0.9) = 2.1

f(1) = 2

f(1.1) = 2.5

f(2) = 5

f(3) = 46

因為最小的f(x) = 2
則arg min f(x) = 1

Reference:

Explanation on arg min

2021/08/31

Spyder: Show Numpy array in variable explorer

By default, Spyder does not show some Numpy arrays in variable explorer.

To show them, simply unselect "Exclude all-uppercase references".

Reference:

why converted numpy array is NOT showing in the variable exloporer in Spyder 4? (StackOverflow)

2021/08/26

Deep Learning: threshold θ and bias b

The output y of a single-node neural network can be represented as:

y = a(Σwixi - θ)

where a is the activation function, xi is the ith input, wis the ith weight, and θ is the threshold.

When the summation of wighted input Σwixi is less than θ, the neuron does not output. Hence we call θ the threshold, which is similar to the threshold potential in a physical neuron.

We may replace θ by -b and get

y = a(Σwixi + b)

where b is called the bias parameter.

So b = - θ is a more generalized representation for the threshold.

References:

A Beginner’s Guide to Neural Networks: Part Two

Hinton Neural Networks課程筆記2b:第一代神經網路之感知機

深度學習的數學:用數學開啟深度學習的大門(博碩) p.12-16

Audacity: Silence the selection of the sound track 將選取的一段音軌靜音

This is a function of Audacity.


To mute a selection in Audacity, simply click the button for 'Silence audio selection'.

Result:

Related Information:

Audacity: Change audio speed and keep the pitch of the talker unchanged (StudyEECC)

Audacity: Change the sample rate from 44.1 kHz to 16 kHz (StudyEECC)