gora.
ai automation · llm integration · remote

AI Automation and LLM Integration Developer

I build LLM-powered features and automation that run in production: Claude and OpenAI API integrations, RAG chatbots over your own content, AI assistants in messengers and workflow automation around the systems you already have.

  • 12 AI models in one production bot
  • Claude, OpenAI, Gemini and Grok in live systems
  • Automated quality-scoring loop on rented GPUs

Concrete kinds of work, each with a system I have run in production.

  • LLM integration into an existing product

    Claude API with prompt caching and streaming over WebSocket, added to an existing NestJS or FastAPI back end without rewriting it.

  • RAG chatbots over your content

    Retrieval over your site or documents with pgvector and voyage-3-large embeddings, answers that cite their sources. chatbot.hiregora.com is my own product built this way: give it a URL and it builds the knowledge base.

  • AI assistants in messengers

    Telegram assistants with memory, files and voice, running several model providers at once. One of mine serves 12 models from four providers in a single chat.

  • Reliability: fallbacks and escalation

    If one provider fails, another one answers. When confidence is low, the question is escalated to a human instead of guessed.

  • AI moderation and classification

    Moderation that learns what a channel is about from its own posts, then classifies each new message as OK, spam, hate or off-topic.

  • Generative media pipelines

    Orchestration of ComfyUI on rented GPUs, LoRA training and automated quality scoring: a model generates, another scores the result on four criteria and the pipeline decides whether to regenerate.

  • Automation with guardrails

    Scheduled pipelines that stop themselves at thresholds and wait for approval, such as a daily catalogue sync that halts if more than a quarter of listings change.

  • Usage tracking and cost control

    Per-request logs, token counters and plan-level limits enforced in the database, so you can see what AI features cost before the invoice arrives.

Three production systems with their numbers. The first two link to full case studies.

GoraGen — AI video generation platform

Multi-tenant platform for AI video: self-serve LoRA training, ComfyUI on rented GPUs and storyboards.

What was done
An iterative quality pipeline: the prompt is translated, Claude picks a model, a rented GPU instance runs ComfyUI, Gemini scores the result on four criteria and the system decides whether to regenerate. A control plane with leader leases recovers after crashes.
Problem solved
Generate usable video without a human reviewing every attempt, while keeping expensive GPU workers from running when nobody needs them.
Technology
  • FastAPI
  • PostgreSQL 16
  • ChromaDB
  • Anthropic Claude
  • Google Gemini 2.5 Pro
  • ComfyUI
  • Vast.ai
  • Docker Compose
In numbers
  • 61 SQL migrations and four apps behind one Caddy
  • Gemini scores each result on 4 criteria before the pipeline decides
  • Idempotency middleware on 5 mutating endpoints

Read the case study →

GoraBot — multi-AI assistant on Telegram

One Telegram chat with 12 models from Anthropic, OpenAI, Google and xAI.

What was done
Chains where four models process a task in sequence, debates with a judge model, project contexts, voice replies and per-plan limits enforced in the database, with request logs for analytics.
Problem solved
Compare and combine different models without losing context, under one rate limit and one billing logic.
Technology
  • python-telegram-bot 21
  • Anthropic, OpenAI, Google, xAI SDKs
  • SQLite
  • FastAPI
In numbers
  • 56 users and 1,783 generations in production on the publication date
  • 5.4M input and 1.06M output tokens, $43.3 in API costs
  • Top models by usage: GPT Thinking 605, Claude Sonnet 492, Claude Opus 311

Read the case study →

chatbot.hiregora.com — RAG chatbot for your site

My own product, in open beta: give it a URL and it builds a knowledge base and answers visitors.

What was done
A crawler (sitemap-aware) builds the knowledge base from a site, Claude answers with retrieval over pgvector and streams the reply over WebSocket, the chat embeds with one script, and the dashboard shows top questions and the topics the bot could not answer.
Problem solved
Give a website an assistant that answers from its own content and cites the sources, without a separate developer or agency.
Technology
  • FastAPI + Celery
  • PostgreSQL 16 + pgvector (HNSW)
  • Redis
  • Claude Haiku 4.5
  • voyage-3-large embeddings
  • Next.js 15 dashboard
  • Preact widget
In numbers
  • About 2 minutes from a URL to the first dialog
  • 1024-dimension voyage-3-large embeddings; answers cite their sources
  • Claude Haiku 4.5 by default, Sonnet 4.6 on the Pro plan

Open the product →

LLM APIs
Anthropic Claude (with prompt caching) · OpenAI GPT · Google Gemini · xAI Grok
Retrieval
pgvector (HNSW) · voyage-3-large embeddings · ChromaDB
Back end
FastAPI · NestJS · PostgreSQL · Redis · WebSocket streaming
Generative media
ComfyUI · kohya_ss LoRA training · Vast.ai GPU rental
Channels
Telegram Bot API · WhatsApp Business API · Web chat widget
Operations
Docker Compose · systemd · Sentry · GitHub Actions

A short path from the problem to something running, with the model's limits spelled out up front.

  1. 01 · First message

    Write to me on Telegram. In your own words, even if you don't know the technical terms. I care more about understanding what you want to build than what you call it.

  2. 02 · Reading between the lines

    Within a day or two I understand the task and give you an estimate of timeline and budget. If I see it isn't my case — I say so right away, no hedging.

  3. 03 · Lock-in

    Price and timeline — in writing. 50% upfront. After that things change only by mutual agreement, and rarely.

  4. 04 · Work

    I show interim results at my own pace. At the end — handover with documentation and access. First month of support is included by default.

  • Direct communication with the developer who writes and runs the system, with no agency layer.
  • I keep a clear line between what the model does and what stays ordinary deterministic code.
  • Production experience with several providers at once, not a single-vendor demo.
  • I plug into existing back ends and products instead of proposing a rewrite.
  • Architecture, back end, front end and deployment from one person.
  • Working remotely with international clients; NDA on request.

I work remotely with companies in different countries, communicating in English or Russian. Most AI projects are done over Telegram, email and video calls.

You keep control of the provider accounts: API keys, billing and data stay with you, and I sign an NDA on request.

Two separate costs apply. Development is scoped per project: what depends on the integration surface, whether retrieval over your documents is involved, how many channels the assistant must cover, and how strict the reliability requirements are. I send a written estimate of timeline and budget, usually within a day or two of your first message.

Running costs go to the model provider and your hosting. For scale: in GoraBot's production database, 1,783 generations across 12 models cost $43.3 in API fees. Cost per request depends heavily on the model, which is why I log tokens per request from the start.

What can AI automate in my business?
Typical candidates are answering repetitive customer questions, triaging and moderating messages, searching internal documents, classifying or scoring content, and preparing drafts for a human to approve. If a task needs judgment that cannot be checked, I will say so and keep a human in the loop.
Claude or OpenAI: which should I use?
I work with both, plus Gemini and Grok. The right choice depends on the task, latency and cost, so I evaluate on your own examples. Several of my systems route between providers and fall back automatically if one is unavailable.
Can you build a RAG chatbot over my website or documents?
Yes. I use pgvector with voyage-3-large embeddings and Claude, and the bot cites the sources it used. chatbot.hiregora.com is my own product built on this: give it a URL and it builds the knowledge base and can be embedded with one script.
Can you add AI to an existing product?
Yes. I add Claude or OpenAI calls to an existing NestJS or FastAPI back end, with streaming and prompt caching where useful, without rewriting the product.
How do you keep AI answers reliable?
With layered checks around the model: rules before a reply is sent, confidence thresholds with escalation to a human, fallbacks between providers, content moderation, and in generative pipelines an automated scoring step that decides whether to regenerate. Nobody can promise a model is never wrong, so these controls limit the damage.
What does it cost to run an AI feature?
Usage is billed by the model provider to your account. As one real data point, GoraBot's 1,783 generations across 12 models used 5.4M input and 1.06M output tokens and cost $43.3. I log tokens per request so you can see costs per feature, and plan limits are enforced in the database.
Do you work with international companies?
Yes. I work remotely with clients in different countries. Communication is in English or Russian, over Telegram, email and video calls.
Who owns the system and the data?
You do. Code goes into your GitHub repository, provider accounts and API keys stay with you, and I hand over documentation and access at the end. I sign an NDA on request and the first month of support is included.

Describe the task in your own words. I respond within 24 hours, usually the same day on weekdays.

Elsewhere: GitHub·Telegram channel·Reddit·X