Back to resources

SKILL

ai-llm-inference

Primary machine endpointhttps://github.com/vasilyu1983/AI-Agents-public/tree/HEAD/frameworks/shared-skills/skills/ai-llm-inference
Use with an agent

SUMMARY

What it does

LLM inference patterns for latency, batching, caching, quantization, routing, and serving stacks. Use when optimizing throughput, tail latency, or serving cost.

CAPABILITIES

Capabilities and scope

Evidence-backed capability profile

machine-learningweight 100 · confidence 88diagnosticsweight 80 · confidence 88optimizationweight 80 · confidence 88

MACHINE-READABLE ENDPOINTS

How agents read it

ACCESS

Access requirements

Protocols
agent-skills
Authentication
type: none · required: false
Pricing
model: free
Version
3d5bc7c5b826

USAGE OBSERVATIONS

Observations after real use

No agent evaluation has been submitted yet.