Published 3 days ago • loading... • Updated 2 days ago
MOREH Showcases High-Performance LLM Inference on AMD GPUs at AMD Advancing AI 2026
Moreh said its MoAI Inference Framework delivered distributed inference on AMD GPUs, highlighting production-ready performance for large language models and agentic AI.
On July 22-23, AMD hosted its Advancing AI 2026 event in San Francisco, where Embedded LLM launched TokenVisor Spaces, a platform providing governed agentic AI services for enterprises and cloud operators on AMD-powered infrastructure.
Agentic AI demands balanced platforms combining CPU-based execution with GPU-accelerated inference, as Dan McNamara, senior vice president and general manager, Compute and Enterprise AI at AMD, noted enterprise AI is moving from experimentation into production.
VAST Data and Embedded LLM collaborated to optimize KV-cache reuse, which Kuntai Du, Chief Scientist and Co-founder at Tensormesh, said is "what makes long-running, multi-turn agents economically viable" for workloads on AMD Instinct GPUs.
Infobell IT Solutions joined the ecosystem, launching four product suites designed to help enterprises modernize infrastructure and manage complex data pipelines and intelligent applications across AMD environments.
TokenVisor Spaces is available for partner deployment, allowing AI cloud operators to sell governed agent services rather than raw GPU hours and providing enterprises auditable workspaces integrated with existing development lifecycles.