On-demand GPU compute, by the hour or by the day

Access dedicated GPUs and cloud AI model APIs on demand. Pay only for what you use while maximizing performance and cost efficiency

Get started

GPU instances

RTX 3090 24GB (RTX 3090)

8 vCPU / 32 GB RAM / 200 GB storage

$4.5/day

$0.22/hr · $120.0/mo

RTX 4090 24GB (RTX 4090)

16 vCPU / 64 GB RAM / 300 GB storage

$7.5/day

$0.35/hr · $200.0/mo

Up to 15% off within 2 days of the start date

RTX A6000 48GB (RTX A6000)

16 vCPU / 64 GB RAM / 300 GB storage

$10.0/day

$0.48/hr · $280.0/mo

A40 48GB (A40)

16 vCPU / 64 GB RAM / 300 GB storage

$9.0/day

$0.4/hr · $240.0/mo

V100 16GB (V100)

8 vCPU / 32 GB RAM / 200 GB storage

$6.0/day

$0.25/hr · $160.0/mo

A100 40GB SXM4 (A100)

32 vCPU / 128 GB RAM / 500 GB storage

$25.0/day

$1.1/hr · $650.0/mo

A100 80GB SXM4 (A100)

32 vCPU / 256 GB RAM / 1000 GB storage

$35.0/day

$1.5/hr · $900.0/mo

Up to 20% off within 3 days of the start date

H100 80GB SXM5 (H100)

32 vCPU / 256 GB RAM / 2000 GB storage

$55.0/day

$2.5/hr · $1400.0/mo

Serverless GPU

Deploy GPU inference endpoints without managing any infrastructure. Highreso Serverless spins up containers on demand, scales to zero when idle, and bills per request or per GPU-second — so you only pay while a request is running. Bring your own container image or pick one of our prebuilt templates for image generation, transcription, or LLM inference to get an endpoint live in minutes.

Image Generation Endpoint

$0.015

per request

Whisper Transcription Endpoint

$0.004

per request

Custom Container (A10G)

$0.0006

per GPU-second

Model Library

Llama-3-8B-Instruct

$0.0003

per 1K tokens

Llama-3-70B-Instruct

$0.0015

per 1K tokens

Mixtral-8x7B-Instruct

$0.0009

per 1K tokens

Stable-Diffusion-XL

$0.012

per request