Pin Abstarct
Pin Abstarct
Abstract

ai on rails · self-hosted models

ai on rails · self-hosted models

Private AI that runs on the server you already own

Private AI that runs on the server you already own

Ruby on Rails software that still works in year five.

Ruby on Rails software that still works in year five.

Small language models installed inside your own infrastructure. No GPU, no new vendor, and your data never leaves your network.

Small language models installed inside your own infrastructure. No GPU, no new vendor, and your data never leaves your network.

CPU no GPU needed

1–2 days to deploy

0 data leaves your network

in production today

A customer categorization model, running on ordinary CPU, inside a background job.

No GPU. No procurement cycle. The model is Gemma, it processes work asynchronously through the existing job queue, and no customer data leaves the network at any point.

Gemma

model in production

Qwen

also CPU-ready

  • Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo
    Client Logo

0-4 yrs

0-4 yrs

average client relationship, in an industry that measures it in months

average client relationship, in an industry that measures it in months

20%

20%

of annual revenue from clients partnering with us for over two years

of annual revenue from clients partnering with us for over two years

5

5

full-time senior Rails engineers in Ahmedabad, plus architecture & DevOps

full-time senior Rails engineers in Ahmedabad, plus architecture & DevOps

100+

100+

products delivered across proptech, fintech, health, IoT and SaaS

products delivered across proptech, fintech, health, IoT and SaaS

the problem with most private ai

It usually starts with a hardware quote

Most private AI conversations start with hardware. A dedicated GPU server, a procurement cycle, lead times, and a spending decision that goes to your finance team before anyone installs anything

For a large share of real business tasks, that is the wrong starting point.

We run a customer categorization model in production today on ordinary CPU, inside a background job, on infrastructure the client already ran. That one is Gemma. It could as easily have been Qwen. Choosing between them for your task and your hardware is part of what we do.

If your task is narrow and your data cannot leave your building, you probably do not need the server everyone is trying to sell you.

why cpu and background jobs

The architecture that makes it affordable

Three things have to be true for this to work. When they are, the hardware problem disappears.

Most business AI does not need to be instant

Categorizing a ticket, tagging a document, extracting invoice fields, routing a lead, summarizing a call note. None needs an answer in 200 milliseconds. It needs an answer before a human next looks at the record.

This is ordinary Rails

Sidekiq or Solid Queue, a worker, retries, a dead letter queue, and a record in Postgres. Your team already understands every part of it. No separate Python service that one person maintains and nobody else can debug.

Small models are good at narrow tasks

Classification, extraction, tagging, routing and short summarization are where a well-prompted small model matches a much larger one. You are not asking it to write your strategy.

A job goes onto the queue, a worker picks it up, the model runs, the result is written back to the database. Your application never blocks. Your users never wait. And a CPU that would be far too slow for a live chat interface is entirely adequate for work that runs in the background.

THE PART PEOPLE GET WRONG

Not on your application server

Most private AI conversations start with hardware. A dedicated GPU server, a procurement cycle, lead times, and a spending decision that goes to your finance team before anyone installs anything

YOUR NETWORK

Rails application

Serves your users. Never blocked.

HTTP

→

Model instance

Gemma or Qwen on CPU. Resize or restart without touching your app.

Nothing crosses this boundary. No prompts, no documents, no results.

Nothing crosses this boundary. No prompts, no documents, no results.

The application gets slower at exactly the moment the model is busiest, which is usually your busiest hour too. So we put the model on its own small instance, and your application calls it over HTTP. The instance is modest and sized to the work.

This is the one cost we will not pretend away. You are not buying a GPU server, but you are running one more small machine. It is a fraction of the alternative, and it is the difference between a model that helps and a model that slows your product down.

This is the one cost we will not pretend away. You are not buying a GPU server, but you are running one more small machine. It is a fraction of the alternative, and it is the difference between a model that helps and a model that slows your product down.

what we install

Five layers, not just a model

Downloading a model takes an afternoon. Making it a feature your business can rely on takes the rest of this list.

01

The model and the runtime

Model selection for your task from the open-weight families that actually run well on CPU, which in practice means Gemma and Qwen. Then a serving setup on your own instance, using a lightweight local runtime rather than a GPU inference server.

02

The application integration

A background job, a queue, retry and failure handling, and the results written into your existing database so your product can use them. This is the part that makes it a working feature rather than a demo.

03

Retrieval over your own documents

When the task needs it, an embedding model and a vector store kept inside your network, so answers are grounded in your content rather than the model's training data.

04

Access control and audit logging

Who is allowed to see what, and a record of every question and answer. In a regulated environment this is not optional, and it is ordinary Rails work.

05

Monitoring and a fallback

What happens when the model is slow, wrong or unavailable. Every AI feature needs a defined behaviour for failure, and most shipped ones do not have it.

why cpu and background jobs

The architecture that makes it affordable

Three things have to be true for this to work. When they are, the hardware problem disappears.

Your data cannot leave your network for legal, contractual or regulatory reasons

Your data cannot leave your network for legal, contractual or regulatory reasons

The task is narrow: classification, tagging, extraction, routing, short summarization

The task is narrow: classification, tagging, extraction, routing, short summarization

The work can run in the background rather than in a live request

The work can run in the background rather than in a live request

You already have infrastructure with spare capacity

You already have infrastructure with spare capacity

You want no per-token bill that grows with your business

You want no per-token bill that grows with your business

You need frontier reasoning. Complex multi-step analysis, long document synthesis and open-ended writing are better served by a large hosted model.

You need instant answers at high volume. A live chat assistant serving thousands of concurrent users needs a GPU. We can size it and install on it, but we will not pretend CPU is enough.

You want someone to buy and rack the servers. We work software only. You provide the machine or the cloud instance.

You have no compliance reason and low volume. Then an API is cheaper and simpler, and we will say so.

scope

Software only, and we mean it

You provide the infrastructure, which for most deployments means one small instance alongside what you already run. We size it with you, select the model, install and configure the runtime, build the integration into your application, set up the retrieval and access control layers where needed, and hand over documentation your team can maintain.

We do not procure hardware, size GPU clusters or resell infrastructure. If your workload genuinely needs a GPU, we will tell you what to buy and work with whoever supplies it.

next step

Tell us what you are trying to classify, extract or route

A Rails engineer will tell you honestly whether a small model on your own infrastructure is the right answer, or whether an API would serve you better. Free, and no obligation either way.

Abstract Image
Abstract Image
Abstract Image

next step

Tell us what you are trying to classify, extract or route

A Rails engineer will tell you honestly whether a small model on your own infrastructure is the right answer, or whether an API would serve you better. Free, and no obligation either way.

Abstract Image
Abstract Image
Abstract Image

next step

Tell us what you are trying to classify, extract or route

A Rails engineer will tell you honestly whether a small model on your own infrastructure is the right answer, or whether an API would serve you better. Free, and no obligation either way.

Abstract Image
Abstract Image
Abstract Image

Testimonial

What Our Clients say

  • Great team, communication and development work. We are very satisfied and can highly recommend working with Essence Solusoft

    We are looking forward to continue working with Sachin and his team in the future. Each developer is very dependable and producing high quality work.

    Beatrice Kern

    Senior Consultant at ThoughtWorks

    I’m so lucky I found Essence!

    I like the fast answers and instantly work with me. Everything is possible with him. Thank you again ?

    Orgad Hayak

    ORGAD International Marketing LTD

    First class Ruby developer

    we worked with Essence Solusoft for 6 months during a transition period of developers internally. Sachin and team were absolutely brilliant from the very beginning. We really appreciated the amount they cared about the app they helped us develop, and the communication and pro-active thinking was just above and beyond.

    Sam Henning

    Co-Founder @ Union Work

    I highly recommend him

    Very fast contact and perfect work. I am very happy with his results! Can only recommend working with him!

    Niklas F

    Blockchain Enthusiast

    Seriously awesome!

    Very fast contact and perfect work. I am very happy with his results! Can only recommend working with him!

    Jason Cosmo

    CEO of Eastern Digital

    I’m really impressed with the quality of his work

    Essence Solusoft is one of the best developers team that I've ever worked with. They are flexible, takes time to understand not just the specs of their work but of the project overall and provide solutions that are synergistic and additive to the entire platform. Essence Solusoft is a pleasure to work with!

    Team Paperwaiter

    A professional, dedicated and collaborative team

    Across engagements, their team has been reliable and supportive, consistently making themselves available when needed, attending scheduled sessions, and working through challenges as they arise. They demonstrate a strong sense of commitment, often putting in the effort required to resolve issues and keep progress moving. They are also approachable and professional to work with - communicative, respectful, and willing to engage in understanding the requirements at hand. In particular, they have a solid working knowledge of OpenProject, which supports their ability to contribute effectively in that space. Overall, they bring a dedicated and collaborative approach to their work, and we value the working relationship.

    Seranne

    CTO at VegaVision Software

    Lead Sr Developer

    Essence Solusoft provided us with resources with which to augment our development team. During our time working with Essence, we were able to complete a migration to a new analytics platform, a project which had been languishing for some time.

    ShirtSpace.com

    CTO at VegaVision Software

Abstract Image
Abstract Image
Abstract Image
Abstract Image

Testimonial

What Our Clients say

Great team, communication and development work. We are very satisfied and can highly recommend working with Essence Solusoft

We are looking forward to continue working with Sachin and his team in the future. Each developer is very dependable and producing high quality work.

Daniel Carter

Co-founder, Union Work (United Kingdom)

I’m so lucky I found Essence!

I like the fast answers and instantly work with me. Everything is possible with him. Thank you again ?

Orgad Hayak

ORGAD International Marketing LTD

First class Ruby developer

we worked with Essence Solusoft for 6 months during a transition period of developers internally. Sachin and team were absolutely brilliant from the very beginning. We really appreciated the amount they cared about the app they helped us develop, and the communication and pro-active thinking was just above and beyond.

Sam Henning

Co-Founder @ Union Work

Testimonial

What Our Clients say

  • Great team, communication and development work. We are very satisfied and can highly recommend working with Essence Solusoft

    We are looking forward to continue working with Sachin and his team in the future. Each developer is very dependable and producing high quality work.

    Beatrice Kern

    Senior Consultant at ThoughtWorks

    I’m so lucky I found Essence!

    I like the fast answers and instantly work with me. Everything is possible with him. Thank you again ?

    Orgad Hayak

    ORGAD International Marketing LTD

    First class Ruby developer

    we worked with Essence Solusoft for 6 months during a transition period of developers internally. Sachin and team were absolutely brilliant from the very beginning. We really appreciated the amount they cared about the app they helped us develop, and the communication and pro-active thinking was just above and beyond.

    Sam Henning

    Co-Founder @ Union Work

    I highly recommend him

    Very fast contact and perfect work. I am very happy with his results! Can only recommend working with him!

    Niklas F

    Blockchain Enthusiast

    Seriously awesome!

    Very fast contact and perfect work. I am very happy with his results! Can only recommend working with him!

    Jason Cosmo

    CEO of Eastern Digital

    I’m really impressed with the quality of his work

    Essence Solusoft is one of the best developers team that I've ever worked with. They are flexible, takes time to understand not just the specs of their work but of the project overall and provide solutions that are synergistic and additive to the entire platform. Essence Solusoft is a pleasure to work with!

    Team Paperwaiter

    A professional, dedicated and collaborative team

    Across engagements, their team has been reliable and supportive, consistently making themselves available when needed, attending scheduled sessions, and working through challenges as they arise. They demonstrate a strong sense of commitment, often putting in the effort required to resolve issues and keep progress moving. They are also approachable and professional to work with - communicative, respectful, and willing to engage in understanding the requirements at hand. In particular, they have a solid working knowledge of OpenProject, which supports their ability to contribute effectively in that space. Overall, they bring a dedicated and collaborative approach to their work, and we value the working relationship.

    Seranne

    CTO at VegaVision Software

    Lead Sr Developer

    Essence Solusoft provided us with resources with which to augment our development team. During our time working with Essence, we were able to complete a migration to a new analytics platform, a project which had been languishing for some time.

    ShirtSpace.com

    CTO at VegaVision Software

Abstract Image
Abstract Image
Abstract Image
Abstract Image

faqs

Frequently Asked Questions

Everything you need to know before getting started with us. Below are out most common questions we get asked.

Can a small model really replace a large one?

For narrow tasks, often yes. Classification, extraction, tagging and routing are jobs where a small model performs comparably to a much larger one. For open-ended reasoning or long-form writing, no, and we will say so rather than sell you a worse result.

Do we need to buy a GPU?

Will this slow down our application?

Where does our data go?

Which models do you work with?

How long does a deployment take?

What does it cost to run?

Direct Access

Tell us about your project.

Tell us about your project.

A senior Rails engineer reads every message received here. Expect a direct response within one working day.

Direct Email

solutions@essencesolusoft.com

Direct WhatsApp / Phone

+91 88665 72265

Studio Location

406, Gravity Retail & Work Spaces, Opp. Sadguru Vatika, Nikol, Ahmedabad 382350, Gujarat, India

Start Your Rails Project

Strict NDA Protected • 100% Confidential