Serving
Preview
Review personalized model usage, deployments, and API access.
Open Serving for personalized model inference. Its tabs are Usage, Models, and How to use. Workspace permissions control whether you can view models, invoke inference, or manage deployments.
Prepare a deployment
- Review the trained version and its evaluation in Models.
- Open Serving's Models tab and inspect the available deployment controls.
- Select the intended version and compatible configuration, then review the resource and cost implications before submitting the deployment.
- Wait for the deployment's ready state before directing traffic to it.
- Use How to use for the exact endpoint and authentication contract.
Use the Usage view to inspect personalized inference activity. Check deployment status and request errors when traffic fails; a model artifact alone does not establish that an endpoint is available.
Serving a trained candidate is separate from the LLM Gateway's hosted catalog. Use the endpoint and model identity shown by the appropriate service rather than assuming the two interfaces are interchangeable. Keep API credentials on the server.