

CPU no GPU needed
1–2 days to deploy
0 data leaves your network
in production today
A customer categorization model, running on ordinary CPU, inside a background job.
No GPU. No procurement cycle. The model is Gemma, it processes work asynchronously through the existing job queue, and no customer data leaves the network at any point.
Gemma
model in production
Qwen
also CPU-ready
0-4 yrs
0-4 yrs

20%
20%
5
5
100+
100+
the problem with most private ai
It usually starts with a hardware quote
Most private AI conversations start with hardware. A dedicated GPU server, a procurement cycle, lead times, and a spending decision that goes to your finance team before anyone installs anything
For a large share of real business tasks, that is the wrong starting point.
We run a customer categorization model in production today on ordinary CPU, inside a background job, on infrastructure the client already ran. That one is Gemma. It could as easily have been Qwen. Choosing between them for your task and your hardware is part of what we do.
If your task is narrow and your data cannot leave your building, you probably do not need the server everyone is trying to sell you.
why cpu and background jobs
The architecture that makes it affordable
Three things have to be true for this to work. When they are, the hardware problem disappears.

Most business AI does not need to be instant
Categorizing a ticket, tagging a document, extracting invoice fields, routing a lead, summarizing a call note. None needs an answer in 200 milliseconds. It needs an answer before a human next looks at the record.

This is ordinary Rails
Sidekiq or Solid Queue, a worker, retries, a dead letter queue, and a record in Postgres. Your team already understands every part of it. No separate Python service that one person maintains and nobody else can debug.

Small models are good at narrow tasks
Classification, extraction, tagging, routing and short summarization are where a well-prompted small model matches a much larger one. You are not asking it to write your strategy.
A job goes onto the queue, a worker picks it up, the model runs, the result is written back to the database. Your application never blocks. Your users never wait. And a CPU that would be far too slow for a live chat interface is entirely adequate for work that runs in the background.
THE PART PEOPLE GET WRONG
Not on your application server
Most private AI conversations start with hardware. A dedicated GPU server, a procurement cycle, lead times, and a spending decision that goes to your finance team before anyone installs anything
YOUR NETWORK
Rails application
Serves your users. Never blocked.
HTTP
→
Model instance
Gemma or Qwen on CPU. Resize or restart without touching your app.
The application gets slower at exactly the moment the model is busiest, which is usually your busiest hour too. So we put the model on its own small instance, and your application calls it over HTTP. The instance is modest and sized to the work.
what we install
Five layers, not just a model
Downloading a model takes an afternoon. Making it a feature your business can rely on takes the rest of this list.
01
The model and the runtime
Model selection for your task from the open-weight families that actually run well on CPU, which in practice means Gemma and Qwen. Then a serving setup on your own instance, using a lightweight local runtime rather than a GPU inference server.
02
The application integration
A background job, a queue, retry and failure handling, and the results written into your existing database so your product can use them. This is the part that makes it a working feature rather than a demo.
03
Retrieval over your own documents
When the task needs it, an embedding model and a vector store kept inside your network, so answers are grounded in your content rather than the model's training data.
04
Access control and audit logging
Who is allowed to see what, and a record of every question and answer. In a regulated environment this is not optional, and it is ordinary Rails work.
05
Monitoring and a fallback
What happens when the model is slow, wrong or unavailable. Every AI feature needs a defined behaviour for failure, and most shipped ones do not have it.
why cpu and background jobs
The architecture that makes it affordable
Three things have to be true for this to work. When they are, the hardware problem disappears.
You need frontier reasoning. Complex multi-step analysis, long document synthesis and open-ended writing are better served by a large hosted model.
You need instant answers at high volume. A live chat assistant serving thousands of concurrent users needs a GPU. We can size it and install on it, but we will not pretend CPU is enough.
You want someone to buy and rack the servers. We work software only. You provide the machine or the cloud instance.
You have no compliance reason and low volume. Then an API is cheaper and simpler, and we will say so.
scope
Software only, and we mean it
You provide the infrastructure, which for most deployments means one small instance alongside what you already run. We size it with you, select the model, install and configure the runtime, build the integration into your application, set up the retrieval and access control layers where needed, and hand over documentation your team can maintain.
We do not procure hardware, size GPU clusters or resell infrastructure. If your workload genuinely needs a GPU, we will tell you what to buy and work with whoever supplies it.
faqs
Frequently Asked Questions
Everything you need to know before getting started with us. Below are out most common questions we get asked.
Can a small model really replace a large one?
For narrow tasks, often yes. Classification, extraction, tagging and routing are jobs where a small model performs comparably to a much larger one. For open-ended reasoning or long-form writing, no, and we will say so rather than sell you a worse result.
Do we need to buy a GPU?
Will this slow down our application?
Where does our data go?
Which models do you work with?
How long does a deployment take?
What does it cost to run?
Direct Access
A senior Rails engineer reads every message received here. Expect a direct response within one working day.
Direct Email
solutions@essencesolusoft.com
Direct WhatsApp / Phone
+91 88665 72265
Studio Location
406, Gravity Retail & Work Spaces, Opp. Sadguru Vatika, Nikol, Ahmedabad 382350, Gujarat, India
Start Your Rails Project























