DeepInfra has officially launched a new AI data center in Toronto, representing the firm’s first overseas deployment outside the U.S. and its ninth facility worldwide. Unveiled on July 8, this new site extends DeepInfra’s global inference infrastructure footprint, amid a widespread industry shift in AI demand from model training to commercial inference workloads.
The Toronto data center features 1.7 MW of power capacity and will house over 1,000 NVIDIA Blackwell B300 GPUs. Upon full operation, the new compute cluster will boost DeepInfra’s global inference scale and deliver reduced network latency for end users across Canada and other global regions.
This expansion aligns with a clear industry-wide transition. As enterprises accelerate the production rollout of generative AI tools, data center infrastructure demand is increasingly focused on inference workloads rather than training tasks. Large language model inference requires consistent GPU availability, low-latency networking, and geographically dispersed infrastructure to power interactive applications, intelligent AI agents, and high-throughput API services. Per McKinsey & Company research referenced in the announcement, inference workloads will exceed 40% of total data center demand by 2030, with a 35% compound annual growth rate.
DeepInfra states the Toronto deployment is purpose-built to meet these evolving needs, placing GPU compute resources closer to end users and local data sources while expanding the platform’s overall throughput. The company independently owns and operates its full GPU infrastructure stack, supporting over 200 open-source AI models. Its managed inference platform processes nearly five trillion tokens weekly, covering both open and proprietary model deployments.
The launch also underscores rising industry adoption of NVIDIA’s Blackwell architecture for modern inference infrastructure. Compared with prior Hopper-generation hardware, Blackwell GPUs deliver enhanced inference throughput and better energy efficiency for large language models and multimodal AI workloads. Though DeepInfra did not release further details on networking and storage architecture, large-scale deployments of this caliber typically leverage high-bandwidth interconnections and distributed storage to maintain full GPU utilization.
As noted by Nikola Borisov, DeepInfra CEO and co-founder, enterprise AI has transitioned from experimental trials to large-scale production rollouts, driving greater demand for globally distributed infrastructure with low latency and scalable inference capabilities. He explained that the Toronto cluster marks the company’s inaugural international expansion beyond U.S. borders, positioning compute resources proximal to global customers and their datasets.
The Toronto facility comes after DeepInfra’s recent Series B financing round and serves as a key part of its long-term infrastructure expansion roadmap. The firm confirms it is evaluating additional international data center sites to meet the growing market demand for GPU-powered inference services.
Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!



