@tilde-nlp/asr-api-client
v5.0.12
Published
Library with services for communicating with Tilde TSP platform
Readme
asr-api-client library
Library is created for implementing dictation. It is written with typescript and also includes typescript models.
Core methods of asrClient
beginVoiceRecognition(resumeLastSession = false) - starts recording (asks permission for mic if needed) and creates connection with asr api. When saveAudio is enabled and resumeLastSession is true, reconnects with the cached session id to resume the previous session; otherwise starts a new session.
endVoiceRecognition(forceCloseWebsocket = false) - stops recording and closes socket if forceCloseWebsocket param is set to true. If forceCloseWebsocket is set to false, it will close connection on next final result.
Audio/journal storage and session resume
Set saveAudio: true in the configuration to have the server persist the incoming audio and a transcript journal for the session. This is the master switch - session resume only works when it is enabled.
When saveAudio is on, ASR responses carry a session_id (GUID). The client caches the latest value (available via asrClient.sessionId) and, when the websocket connection drops and is retried (see retryIfInteruppted), automatically reconnects with the cached sessionId so the same session continues (same audio file, continuous journal).
Things to be aware of:
- Resume can be refused. The server only lets a reconnect claim the session once it knows the previous connection is gone - i.e. the session was marked suspended or its heartbeat went stale (liveness timeout, typically ~30s). If you reconnect before that, or the audio format (
sampleRate/channelCount) differs, or the session has expired, the server silently starts a brand-new session instead. ConfigurewsConnectionRetryDelayto exceed the server's liveness timeout - the default (iteration * 1000) retries after 1s, which will be refused and orphan the original session. Use theonSessionChanged(newSessionId, previousSessionId)callback to detect a refused resume - a definedpreviousSessionIdthat differs fromnewSessionIdmeans the resume did not take. - No audio replay. The server does not buffer or replay unsent audio. Anything not sent before the drop is absent from the stored file and transcript. If you need gapless audio, buffer unacknowledged chunks in your app and resend them after reconnect.
pause()/resume()(orsendControlMarker) record journal markers only - they do not affect recognition or the connection.- The cached session id survives EOS and a clean stop (
endVoiceRecognition). On the next start you decide:beginVoiceRecognition(true)reconnects with the cached id to resume the last session,beginVoiceRecognition()(default) discards it and starts a new session.
Real time oscilogram
Library also provides possibility to draw real time oscilogram that will be drawn into <canvas> element. To be able to create this visualization, you need to add <canvas></canvas> element to your html and pass elements ID to visualizer configuration.
Examples
Couple small examples for easier integration. Remember to replace variable placeholders with your own values.
Plain js without any visualization
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta http-equiv="X-UA-Compatible" content="IE=edge">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Tilde speech recognition example</title>
<!-- Include Tilde asr api client -->
<script src="https://unpkg.com/@tilde-nlp/asr-api-client@latest/index.js"></script>
</head>
<body>
<script language="javascript">
// set up config
var config = {
url: "wss://services.tilde.com/service/asr/ws/${SYSTEM_NAME}/?contentType=audio/x-raw&sampleRate=44100&channelCount=1&x-api-key=${API_KEY}",
onResult: result => { if (result.final) console.log(result) }, // partial or final result
onRecordingStartStop: isRecording => console.log(isRecording), // boolean value emitted whenever isRecording changes
onError: error => console.error(error) // error callback
}
// Create asr client
var asrClient = new window["asr-api-client"].AsrClient(config);
// Start voice recognition
asrClient.beginVoiceRecognition();
</script>
</body>
</html>Plain js with audio visualization
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta http-equiv="X-UA-Compatible" content="IE=edge">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Tilde speech recognition example</title>
<!-- Include Tilde asr api client -->
<script src="https://unpkg.com/@tilde-nlp/asr-api-client@latest/index.js"></script>
</head>
<body>
<script language="javascript">
// set up config
var config = {
url: "wss://services.tilde.com/service/asr/ws/${SYSTEM_NAME}/?contentType=audio/x-raw&sampleRate=44100&channelCount=1&x-api-key=${API_KEY}",
onResult: result => { if (result.final) console.log(result) }, // partial or final result
onRecordingStartStop: (isRecording, ctx) => {
// start drawing audio oscilogram if neccesarry
if (isRecording) {
ctx.audioVisualizer?.visualizeAudio();
}
}, // boolean value emitted whenever isRecording changes
onError: error => console.error(error), // error callback
// configuration if you want to add real time oscillogram, if not, you do not have to set this value
visualizerConfig: {
// id of container where to put visualizer
visualizerId: "audio-visualizer",
// optional param - visualization color. Default value: "#811331"
strokeStyle: "#811331"
}
}
// Create asr client
var asrClient = new window["asr-api-client"].AsrClient(config);
// Start voice recognition
asrClient.beginVoiceRecognition();
</script>
</body>
<!--necessary if you need real time oscilogram-->
<canvas id="audio-visualizer"></canvas>
</html>