Use cases
AI compute for inference

LLM Inference GPU Workspaces

Serve models on GPU-backed workspaces for prototypes, internal tools, API testing, and low-latency inference.

Ready template

Recommended workspace

Template

Ollama

GPU

H100

Access

Browser + SSH

Billing

Per-hour

porfal create --use-case llm-inferenceworkspace ready: browser + ssh

Workloads

Ollama endpoints

vLLM servers

Open WebUI

Private model demos

Recommended stack

Ollama

vLLM

Ubuntu CUDA

Docker

GPU fit

H100

H200

RTX 6000 Ada

Why Porfal

Dedicated GPU compute without infrastructure work.

Pick a GPU, launch the template, connect from browser or SSH, and stop the workspace when the run is done.

Run private inference without managing GPU hosts.

Use browser and SSH access for debugging.

Pay per running workspace while testing models.

Workflow

A short path from GPU to running workload.

1

Choose VRAM

2

Select Ollama or vLLM

3

Load model

4

Connect API or browser