copilotedge-ai
v1.0.0
Published
The evolution of copilotedge. One-line CopilotKit + Cloudflare AI integration with built-in telemetry, cost tracking, and 93% savings.
Maintainers
Readme
copilotedge-ai
🚀 The evolution of copilotedge (10,000+ downloads). One-line CopilotKit + Cloudflare AI integration with 93-99.7% cost savings vs GPT-5.
October 2025 Update: With OpenAI's GPT-5 now live at $1.25-10/M tokens, Google's Gemini 2.5, Claude Opus 4.1, and other frontier models raising the bar on capabilities and costs, Cloudflare AI through copilotedge-ai is more valuable than ever. Get frontier-like performance for a fraction of the price.
Quick Start
Installation
npm install copilotedge-ai
# or
pnpm add copilotedge-aiBasic Usage
// app/api/copilotedge/route.ts
import copilotedge from "copilotedge-ai";
export const { POST, OPTIONS } = copilotedge();Use with CopilotKit
// app/layout.tsx
import { CopilotKit } from "@copilotkit/react-core";
import "@copilotkit/react-ui/styles.css";
export default function RootLayout({ children }) {
return (
<html lang="en">
<body>
<CopilotKit runtimeUrl="/api/copilotedge">{children}</CopilotKit>
</body>
</html>
);
}Configuration Options
Choose Your Model
export const { POST } = copilotedge({
model: "SMART", // or 'FAST', 'CODE', 'CREATIVE', 'VISION'
});Available models:
FAST- Llama 3.1 8B (default) - Fastest responsesSMART- Llama 3.3 70B - Best reasoningCODE- DeepSeek Coder - Optimized for programmingCREATIVE- Llama 3.1 70B - Creative writingVISION- LLaVA 1.5 - Multimodal (text + images)
Enable Telemetry
Track usage, costs, and performance:
export const { POST } = copilotedge({
telemetry: true, // Simple telemetry to console
});
// Or with custom callback
export const { POST } = copilotedge({
telemetry: {
enabled: true,
callback: async (data) => {
console.log("Tokens:", data.tokens.total);
console.log("Cost: $", data.cost.total.toFixed(6));
console.log("Latency:", data.latency, "ms");
// Send to your analytics
await analytics.track("ai-usage", data);
},
userId: "user-123",
sessionId: "session-456",
},
});Telemetry data includes:
- Token counts - Input, output, and total
- Costs - Calculated per request in USD
- Latency - Response time in milliseconds
- Model - Which model was used
- User/Session IDs - For tracking
Presets for Common Use Cases
import { presets } from "copilotedge-ai";
// For chat applications
export const { POST } = presets.chat();
// For code assistants
export const { POST } = presets.code();
// For creative writing
export const { POST } = presets.creative();
// For customer support
export const { POST } = presets.support();Advanced Configuration
export const { POST } = copilotedge({
// Model selection
model: "SMART",
// Fine-tuning
temperature: 0.7, // 0-1, higher = more creative
maxTokens: 2000, // Maximum response length
// System prompt
systemPrompt: "You are a helpful AI assistant.",
// Telemetry
telemetry: true,
// Debugging
debug: true, // Logs requests and responses
// Override credentials (optional)
accountId: process.env.MY_CF_ACCOUNT,
apiToken: process.env.MY_CF_TOKEN,
});Environment Variables
Set these in your .env.local:
CLOUDFLARE_ACCOUNT_ID=your-account-id
CLOUDFLARE_API_TOKEN=your-api-tokenGetting Your Cloudflare Credentials
Sign up/Login at dash.cloudflare.com
Get your Account ID:
- Find it in the right sidebar of your dashboard
- Or in the URL:
dash.cloudflare.com/[ACCOUNT_ID]/...
Create an API Token:
- Click your profile icon (top right) → "My Profile"
- Go to "API Tokens" tab
- Click "Create Token" → "Custom token"
- Configure with these permissions:
- Account → Workers AI → Read
- Account → Workers AI → Edit
- Under "Account Resources" → Include → Select your account
- Click "Continue to summary" → "Create Token"
- Copy the token (you won't see it again!)
Verify it works (optional):
curl -X GET "https://api.cloudflare.com/client/v4/accounts/YOUR_ACCOUNT_ID/ai/models/search" \ -H "Authorization: Bearer YOUR_API_TOKEN"
Cost Comparison (Updated October 2025)
| Provider | Input Cost (per 1M tokens) | Output Cost (per 1M tokens) | Total (1M tokens) | | ----------------------------------- | -------------------------- | --------------------------- | ----------------- | | OpenAI GPT-5 (released Aug 2025) | $1.25 | $10.00 | $11.25 | | OpenAI GPT-5 Mini | $0.25 | $2.00 | $2.25 | | OpenAI GPT-5 Nano | $0.05 | $0.40 | $0.45 | | Claude Opus 4.1 / Sonnet 4.5 | ~$3-15 | ~$15-75 | ~$18-90 | | Google Gemini 2.5 Pro | ~$1.25-7 | ~$5-21 | ~$6.25-28 | | Cloudflare (via copilotedge-ai) | $0.01 | $0.02 | $0.03 |
Save 93-99.7% on AI costs compared to frontier models!
What About GPT-5?
OpenAI released GPT-5 in August 2025 with impressive capabilities (272K context, agentic workflows, multimodal), but even their most economical variant (GPT-5 Nano at $0.45/1M tokens) is 15x more expensive than Cloudflare AI.
For most applications (chat, support, content generation, code assistance), Cloudflare's Llama 3.1/3.3 models deliver comparable quality at a fraction of the cost.
Performance
Average response times:
- Cloudflare: 110ms ⚡
- OpenAI: 350ms
- Anthropic: 420ms
Migration from copilotedge (deprecated)
If you're using the original copilotedge package:
- import { createCopilotEdgeHandler } from 'copilotedge';
+ import copilotedge from 'copilotedge-ai';
- export const POST = createCopilotEdgeHandler({
- apiKey: process.env.CLOUDFLARE_API_TOKEN,
- accountId: process.env.CLOUDFLARE_ACCOUNT_ID,
- });
+ export const { POST } = copilotedge();Examples
Next.js App Router
// app/api/copilot/route.ts
import copilotedge from "copilotedge-ai";
export const { POST, OPTIONS } = copilotedge({
model: "SMART",
telemetry: true,
});Next.js Pages Router
// pages/api/copilot.ts
import copilotedge from "copilotedge-ai";
const { POST } = copilotedge();
export default async function handler(req, res) {
if (req.method === "POST") {
return POST(req);
}
res.status(405).end();
}With Custom Analytics
import copilotedge from "copilotedge-ai";
import { analytics } from "@/lib/analytics";
export const { POST } = copilotedge({
telemetry: {
enabled: true,
callback: async (data) => {
// Track in your analytics
await analytics.track("ai.request", {
...data,
timestamp: new Date().toISOString(),
});
// Alert on high costs
if (data.cost.total > 0.001) {
console.warn("High cost request:", data);
}
},
},
});The 2025 AI Landscape
As of October 2025, the AI model ecosystem has evolved significantly:
Frontier Models Now Live
- GPT-5 (OpenAI, Aug 2025): The flagship model with 272K context, deep reasoning, agentic workflows, and multimodal support. Three variants: full ($1.25-10/M), mini ($0.25-2/M), nano ($0.05-0.40/M).
- Claude Opus 4.1 & Sonnet 4.5 (Anthropic): Safety-focused, excellent for code and agentic tasks. Premium pricing (~$3-75/M tokens).
- Gemini 2.5 Pro (Google): Fully multimodal (text/image/audio/video), "thinking mode" for reasoning. Tightly integrated with Google's ecosystem.
- Qwen3 / Qwen2.5-VL (Alibaba): Open-licensed (Apache 2.0), competitive quality, multimodal, 128K+ context.
- Mistral Magistral (Mistral AI): Open weights, reasoning-capable, edge-deployable.
Key Trends
- Agentic workflows - Models can now chain reasoning, call APIs, browse, execute code, maintain memory
- Longer contexts - 200K-400K tokens becoming standard
- Multimodal everywhere - Text+image+audio+video in single models
- Safety & alignment - Red-teaming, guardrails, auditability now standard
- Open weight revival - More customizable, fine-tunable, on-prem options
- Cost commoditization - Even frontier models are dropping prices
Why Cloudflare AI Still Wins
Despite these advances, Cloudflare AI remains the most economical choice for 90% of use cases:
- ✅ Chat applications - Llama 3.3 70B matches GPT-5 quality for conversation
- ✅ Customer support - Fast responses, low latency, fraction of the cost
- ✅ Content generation - Creative writing, summaries, translations
- ✅ Code assistance - DeepSeek Coder rivals specialized models
- ✅ Prototyping & MVPs - Ship fast without budget concerns
When to use premium models:
- Multi-step reasoning over 100K+ tokens
- Mission-critical medical/legal analysis
- Advanced multimodal tasks (video understanding)
- Cutting-edge research requiring latest capabilities
For everything else? copilotedge-ai gives you 95% of the capability at 1% of the cost.
API Reference
copilotedge(options?)
Creates Next.js route handlers for CopilotKit.
Options:
model- Model preset or name (default: 'FAST')telemetry- Enable telemetry (boolean or config object)temperature- Model temperature 0-1 (default: 0.7)maxTokens- Max response tokens (default: 1000)systemPrompt- System instructionsdebug- Enable debug loggingaccountId- Override Cloudflare account IDapiToken- Override Cloudflare API token
Returns:
POST- Next.js POST route handlerOPTIONS- Next.js OPTIONS route handler for CORS
presets
Pre-configured setups for common use cases:
presets.chat()- Optimized for chat applicationspresets.code()- Optimized for code generationpresets.creative()- Optimized for creative writingpresets.support()- Optimized for customer support
Resources (work in progress)
ChatCloudflareWorkersAI
https://js.langchain.com/docs/integrations/chat/cloudflare_workersai/
Workers AI allows you to run machine learning models, on the Cloudflare network, from your own code.
This will help you getting started with Cloudflare Workers AI chat models. For detailed documentation of all ChatCloudflareWorkersAI features and configurations head to the API reference.
https://www.npmjs.com/package/@langchain/cloudflare
https://js.langchain.com/docs/concepts/chat_models/
