Use cases
AI compute for inference
LLM Inference GPU Workspaces
Serve models on GPU-backed workspaces for prototypes, internal tools, API testing, and low-latency inference.
Ready template
Recommended workspace
Template
Ollama
GPU
H100
Access
Browser + SSH
Billing
Per-hour
porfal create --use-case llm-inferenceworkspace ready: browser + sshWorkloads
Ollama endpoints
vLLM servers
Open WebUI
Private model demos
Recommended stack
Ollama
vLLM
Ubuntu CUDA
Docker
GPU fit
H100
H200
RTX 6000 Ada
Why Porfal
Dedicated GPU compute without infrastructure work.
Pick a GPU, launch the template, connect from browser or SSH, and stop the workspace when the run is done.
Run private inference without managing GPU hosts.
Use browser and SSH access for debugging.
Pay per running workspace while testing models.
Workflow
A short path from GPU to running workload.
1
Choose VRAM
2
Select Ollama or vLLM
3
Load model
4
Connect API or browser