GT-ASSIST-V2
One Kenyan account for multi-model AI chat, image, video and voice — billed in KES.
The problem
A Kenyan team that wants AI in its daily work ends up juggling half a dozen separate tools — one for chat, one for images, one for video, one for transcription — each with its own USD subscription, its own login, and no way to pay in shillings. Usage is impossible to budget, and most tools cannot even be trialled without a foreign card. GT-ASSIST-V2 puts the whole set behind a single account with one credit balance and local payment rails.
What the agent does
- 1The user signs in and receives a daily free credit allowance; top-ups go through a Paystack checkout they confirm themselves.
- 2They send a chat message; the app routes the request to a model suited to the job.
- 3The reply streams back token by token, with tool calls (web search, knowledge base) surfaced inline.
- 4Media requests branch off the same credit balance: images, image edits, video animation, speech-to-text and text-to-speech.
- 5Conversations, media and the credit ledger sync to the user's account so any device resumes where they left off.
- 6Rate-limited provider calls retry with exponential backoff and jitter rather than failing the user's request.
Agent architecture
Models
- DeepSeek V3.2 / V4-Flash (fast everyday chat, 1 credit)
- Kimi K2.6 / K3 (long-context work)
- GLM-5.1 / GLM-5.2 and Qwen3 235B-Thinking (hard reasoning lane)
- Qwen-Image-2512 + FLUX.1 schnell (images), LTX 2.5 (i2v/t2v video), Whisper (STT), Kokoro TTS
Tools
- Chutes API (chat, image, video, audio)
- Clerk (identity) and Supabase (cloud sync, edge functions)
- Paystack (KES payments, webhook confirmation)
- Web search + knowledge-base grounding (searchAgent / knowledgeBase)
- retryFetch (429/503 retry with exponential backoff + jitter)
Memory
Supabase persists conversations, generated media and the per-user credit ledger across devices; the conversation branch/edit history is part of that state.
Routing
The user picks a model per task, and the platform groups them by cost and depth: 1-credit lanes (DeepSeek V4-Flash, Qwen3.6-27B) for everyday chat, reasoning lanes (Kimi K3, GLM-5.2, Qwen3-235B-Thinking) for hard work, with media models routed by task type.
Hosting
Tier 1 checklist
5 of 12 passingEvery project must clear all 12 Tier 1 checks before it is presented as live. This is enforced in code, not by convention.
- README explains the architecture in ≤ 5 minutesMissing — README has no architecture section yet.
- 60-second demo video recordedMissing — 60-second walkthrough not recorded.
- Public GitHub repositoryMissing — repo is private pending a user call and a secret sweep.
- Prompts versioned
- Every model call loggedUnverified — no per-call log surfaced yet.
- Tool outputs schema-validatedUnverified — tool output schemas not demonstrated.
- Retries with backoff
- Secrets / PII stripped before the modelFAILING — the live bundle ships a Chutes API key inlined by Vite. Fix: rotate the key and proxy Chutes calls server-side.
- Human approval on money, email and delete actions
- Streaming responses
- Repeat queries cachedUnverified — no repeat-query cache demonstrated.
- Token / time / cost budget cap
Not presented as live yet — 7 of 12 Tier 1 checks outstanding.