Platform
What VijayCloud actually gives you
A private GPU cloud on our own hardware in Mumbai, a control plane that provisions and watches it, and Medusa, our private AI assistant, running on the same infrastructure. The configurations below are what we provision today.
- GPUs
- NVIDIA A100 · H100
- Fabric
- InfiniBand 200 Gb/s
- Encryption
- AES-256, at rest and in transit
- Time to live
- Under 24 hours
Services
Three things we run for you. Everything else on this page is a property of one of them.
Dedicated GPU instances
Single-tenant A100 and H100 nodes reserved to you. You get the whole GPU, not a slice of one, so throughput does not move when someone else starts a job.
Managed provisioning
Tell us the workload, the GPU requirement and the budget. We size, build and tune the environment rather than handing you an empty console. Typically live in under 24 hours.
Medusa, private AI assistant
Several open models served from our own infrastructure, so prompts and outputs never leave your tenancy. Try Medusa.
Features
The core platform, in the order that tends to matter when you are deciding.
Dedicated GPU power, no sharing
GPUs are reserved to a single tenant. No noisy neighbours, and no variance in training times you cannot explain.
Fully customisable environments
Choose the GPUs, frameworks and configuration per project. Up to 8 x A100 with 512 GB memory and 4 TB NVMe in a single environment.
Private and secure by design
Isolated single-tenant environments, AES-256 encryption at rest and in transit, and a private network. Your data and models stay yours.
Usage-based pricing
Pay for the resources you use. No bundled extras and no month-end surprises. See pricing.
Fast deployment
Under 24 hours from sign-up to a running environment, with a person to talk to rather than a queue.
Scales with the work
Start on a single GPU and grow. The control plane autoscaler handles capacity as jobs change shape.
AI automation
Two layers: the control plane that runs the infrastructure, and Medusa running on top of it.
Medusa serves several open models from our own hardware. You pick the model per conversation:
- gemma3:12b — default
- gemma4:12b
- qwen3.6
- deepseek-r1:14b
Because the models run on our infrastructure rather than a third-party API, nothing you type is sent to an external provider or used to train someone else’s model.
Monitoring and notifications
Every environment reports its own state. The dashboard is the primary surface for it.
Routing events elsewhere. Events are surfaced in the dashboard. If you need them pushed into email or a chat channel your team already watches, raise it while we are provisioning — environments are configured per tenancy, so it is a setup question rather than a feature request. Ask us.
Real-time analytics
Live telemetry per GPU, not a daily rollup. This is the same data the autoscaler and billing meter read.
Utilisation
GPU utilisation as a percentage, plotted over the last hour, per device.
Thermals and power
Die temperature and power draw in watts — the two numbers that tell you whether a node is throttling before throughput drops.
System resources
vCPU, memory and NVMe consumption alongside GPU load, so you can see which one is actually the bottleneck.
Usage and spend
Metered consumption feeding your invoice, visible as it accrues rather than at month end. How billing works.
Integrations
Your stack, not ours. Environments ship with the frameworks you name at provisioning time.
- PyTorch
- TensorFlow
- Hugging Face
- Jupyter
- Docker
Something not listed? Environments are built per tenancy rather than from a fixed catalogue, so an unusual framework or driver version is usually a configuration detail. Tell us what you need.
Not sure which configuration fits?
Describe the workload and we will size it. Your first month carries a 14-day money-back guarantee, so getting it slightly wrong is not expensive.