@assemblyline-agents/audio
v10.4.0
Published
Pre-model audio processing adapters for Assembly Line agents.
Downloads
8,020
Maintainers
Readme
@assemblyline-agents/audio
Reusable, channel-scoped pre-model audio processing for Assembly Line agents.
Attachment-capable channel profiles select this package automatically. Photon, Slack, Telegram, and Teams use OpenRouter transcription by default:
channels: [photon]Override or disable audio inside the channel:
channels:
photon:
audio: openaiSet OPENROUTER_API_KEY for the default profile. Its default model is
openai/whisper-large-v3-turbo; override it with
OPENROUTER_AUDIO_TRANSCRIPTION_MODEL or channel options. The OpenAI profile
uses OPENAI_API_KEY and defaults to gpt-4o-mini-transcribe; override it with
OPENAI_AUDIO_TRANSCRIPTION_MODEL.
Successful transcripts are stored in the configured private blob adapter and reused by attachment hash. Transcript text is supplied to the model as untrusted derived context and is not copied into run audit events.
Apple CAF voice notes are converted to M4A with the package's bounded FFmpeg binary before they are sent to the selected transcription provider. The original private attachment is not modified. M4A and other provider-supported formats are sent unchanged.
