Image

This Workstation Is Actually a Call Centre

At first glance, this might seem like any other workstation, but this is a fully functional call centre running under the hood.

Here is a look at what it takes to build a localized AI-driven call centre from the ground up.

The Challenge

Our client is one of India’s largest e-commerce shipping and logistics platforms. Behind every order it moves sits a phone call someone has to make: confirming the order, coordinating the delivery, reminding a buyer about a payment, booking an appointment.

The team is building AI voice agents to handle those calls at high volume. The brief to us was short: run every model and every agent fully on premises, on hardware they own.

An AI phone call looks like a simple one way conversation.But behind the scenes, it is a relay between several models and agents running in tandem.

  • Listening. A speech recognition model turns the caller’s voice into text.
  • Understanding. A language model works out what the caller wants.
  • Doing. An AI agent acts on it: it pulls the order record from a database, checks a delivery slot, logs a payment promise, or books the appointment.
  • Speaking. A text-to-speech model turns the reply back into a natural voice.

Now this is one call, A call centre runs hundreds of them at once. While one agent waits on a database lookup, others are mid-sentence, and every one of them keeps sending fresh requests to the models.

So the machine has two very different workloads running side by side:

  • The model layer. Speech recognition, language, voice generation and embedding models, often loaded at the same time. This work lives on the GPUs, and it is hungry for VRAM.
  • The agent layer. Every live conversation, its state, its database calls, its retrieval lookups, and the job of keeping the GPUs fed. This work lives on the CPU, and it is hungry for cores.

The Solution

The model layer: 2x NVIDIA RTX PRO 6000 Blackwell

Each RTX PRO 6000 Blackwell Workstation Edition carries 96GB of GDDR7 with ECC, so the system holds 192GB of GPU memory across two cards. That is two separate 96GB pools, and the software decides what lives where.

For a voice agent stack, VRAM is the constraint that matters most. Speech, language, voice and embedding models may all need to stay loaded at once. Two cards give the client three ways to use them: split one large model across both, pin different models to different GPUs, or run several copies of the same model in parallel to serve more calls.

The agent layer: AMD Ryzen Threadripper PRO 9985WX

The 9985WX brings 64 cores and 128 threads, enough to hold many live conversations, databases and retrieval systems at once while still feeding both GPUs without a queue forming at the CPU.

Memory: 384GB DDR5 ECC RDIMM

Four 96GB registered ECC modules at 5,200 MT/s, one per memory channel on the platform. There’s also a lot of room to expand as only half the slots are populated.

Why On Premises

Think about how many calls a shipping platform this size makes in a day. Now put every one of them on rented GPUs. The cloud bill gets big fast, and it grows again every time the team tries a newer model, because experiments burn GPU hours too.

Owning the machine flips that: test as much as you like, the hardware is already paid for. There is an added benefit of privacy as well. These calls handle payment details, and those details stay inside the client’s own walls, on a machine the client owns outright.

No more cloud bills.

The client now develops and runs its AI voice agents on hardware it owns. Order confirmations, delivery calls, payment reminders and appointment bookings no longer add GPU hours to a cloud invoice. Every call runs through one workstation, models and agents together, and the cost was paid once.

Owning the Compute

More companies and startups are looking at their cloud bills and asking one question: at what point does owning the hardware beat renting it? We help you answer that for your own workload, then build the machine that fits it.

If you want to see what a setup like this looks like for your use case, book a spec call with our team.

SHARE THIS POST

Leave a Reply

Your email address will not be published. Required fields are marked *