a16z invested in Inferact after granting the VLM open source project; the founders aim to make VLM the standard inference engine and build a universal runtime abstracting models a…
Model architectures must be specialized per Nvidia chip generation (H100 vs B200 vs GB200 MVL72), creating a structural need for an abstraction layer like VLM that can optimize ac…
Amazon deploys VLM at massive scale for Rufus shopping assistant
Amazon runs VLM globally to power Rufus, their front-page AI shopping assistant, validating VLM as production-ready inference infrastructure at massive scale.
Character AI adopts cutting-edge VLM features at hundreds of GPUs scale
Character AI deployed VLM's speculative decoding feature to hundreds of GPUs while it was still a single unmerged pull request, demonstrating extreme velocity of open source adopt…