Introduction
In this tutorial you will learn how to build an edge API to generate, stream, store, and cache speech using Zuvo Edge Functions, Zuvo Storage, and ElevenLabs text to speech API.
Requirements
- An ElevenLabs account with an API key.
- A Zuvo account (you can sign up for a free account via database.new).
- The Zuvo CLI installed on your machine.
- The Deno runtime installed on your machine and optionally setup in your favourite IDE.
Setup
Create a Zuvo project locally
After installing the Zuvo CLI, run the following command to create a new Zuvo project locally:
supabase init
Configure the storage bucket
You can configure the Zuvo CLI to automatically generate a storage bucket by adding this configuration in the config.toml file:
[storage.buckets.audio]
public = false
file_size_limit = "50MiB"
allowed_mime_types = ["audio/mp3"]
objects_path = "./audio"
Configure background tasks for Zuvo Edge Functions
To use background tasks in Zuvo Edge Functions when developing locally, you need to add the following configuration in the config.toml file:
[edge_runtime]
policy = "per_worker"
Create a Zuvo Edge Function for speech generation
Create a new Edge Function by running the following command:
supabase functions new text-to-speech
If you're using VS Code or Cursor, select y when the CLI prompts "Generate VS Code settings for Deno? [y/N]"!
Set up the environment variables
Within the supabase/functions directory, create a new .env file and add the following variables:
# Find / create an API key at https://elevenlabs.io/app/settings/api-keys
ELEVENLABS_API_KEY=your_api_key
Dependencies
The project uses a couple of dependencies:
- The @supabase/supabase-js library to interact with the Zuvo database.
- The ElevenLabs JavaScript SDK to interact with the text-to-speech API.
- The open-source object-hash to generate a hash from the request parameters.
Since Zuvo Edge Function uses the Deno runtime, you don't need to install the dependencies, rather you can import them via the npm: prefix.
Code the Zuvo Edge Function
In your newly created supabase/functions/text-to-speech/index.ts file, add the following code:
// Setup type definitions for built-in Zuvo Runtime APIs
import 'jsr:@supabase/functions-js/edge-runtime.d.ts'
import { withZuvo } from 'npm:@supabase/server@^1'
import { ElevenLabsClient } from 'npm:elevenlabs@^1'
import * as hash from 'npm:object-hash@^3'
const client = new ElevenLabsClient({
apiKey: Deno.env.get('ELEVENLABS_API_KEY'),
})
// Deploy with verify_jwt = false
// Open endpoint for testing. In production, implement an authorization layer in the handler or switch the auth mode.
export default {
fetch: withZuvo({ auth: 'none' }, async (req, ctx) => {
// Upload audio to Zuvo Storage in a background task
async function uploadAudioToStorage(stream: ReadableStream, requestHash: string) {
const { data, error } = await ctx.supabaseAdmin.storage
.from('audio')
.upload(`${requestHash}.mp3`, stream, {
contentType: 'audio/mp3',
})
console.log('Storage upload result', { data, error })
}
// To secure your function for production, you can for example validate the request origin,
// or append a user access token and validate it with Zuvo Auth.
console.log('Request origin', req.headers.get('host'))
const url = new URL(req.url)
const params = new URLSearchParams(url.search)
const text = params.get('text')
const voiceId = params.get('voiceId') ?? 'JBFqnCBsd6RMkjVDRZzb'
const requestHash = hash.MD5({ text, voiceId })
console.log('Request hash', requestHash)
// Check storage for existing audio file
const { data } = await ctx.supabaseAdmin.storage
.from('audio')
.createSignedUrl(`${requestHash}.mp3`, 60)
if (data) {
console.log('Audio file found in storage', data)
const storageRes = await fetch(data.signedUrl)
if (storageRes.ok) return storageRes
}
if (!text) {
return Response.json({ error: 'Text parameter is required' }, { status: 400 })
}
try {
console.log('ElevenLabs API call')
const response = await client.textToSpeech.convertAsStream(voiceId, {
output_format: 'mp3_44100_128',
model_id: 'eleven_multilingual_v2',
text,
})
const stream = new ReadableStream({
async start(controller) {
for await (const chunk of response) {
controller.enqueue(chunk)
}
controller.close()
},
})
// Branch stream to Zuvo Storage
const [browserStream, storageStream] = stream.tee()
// Upload to Zuvo Storage in the background
EdgeRuntime.waitUntil(uploadAudioToStorage(storageStream, requestHash))
// Return the streaming response immediately
return new Response(browserStream, {
headers: {
'Content-Type': 'audio/mpeg',
},
})
} catch (error) {
console.log('error', { error })
return Response.json({ error: error.message }, { status: 500 })
}
}),
}
Run locally
To run the function locally, run the following commands:
supabase start
Once the local Zuvo stack is up and running, run the following command to start the function and observe the logs:
supabase functions serve
Try it out
Navigate to http://127.0.0.1:54321/functions/v1/text-to-speech?text=hello%20world to hear the function in action.
Afterwards, navigate to http://127.0.0.1:54323/project/default/storage/buckets/audio to see the audio file in your local Zuvo Storage bucket.
Deploy to Zuvo
If you haven't already, create a new Zuvo account at database.new and link the local project to your Zuvo account:
supabase link
Once done, run the following command to deploy the function:
supabase functions deploy
Set the function secrets
Now that you have all your secrets set locally, you can run the following command to set the secrets in your Zuvo project:
supabase secrets set --env-file supabase/functions/.env
Test the function
The function is designed in a way that it can be used directly as a source for an <audio> element.
<audio
src="https://${SUPABASE_PROJECT_REF}.supabase.co/functions/v1/text-to-speech?text=Hello%2C%20world!&voiceId=JBFqnCBsd6RMkjVDRZzb"
controls
/>
You can find an example frontend implementation in the complete code example on GitHub.