San Francisco-based Together AI has signed a $240 million multi-year agreement with IBM to deploy NVIDIA AI infrastructure on IBM Cloud. The arrangement will support large-scale AI inference and expand Together AI’s capacity to provide open-source models and AI services to enterprise customers.
Key Takeaways
- Together AI signed a $240 million multi-year agreement with IBM.
- The San Francisco company will use NVIDIA HGX B300 infrastructure on IBM Cloud.
- The infrastructure is designed to support large-scale AI inference.
- Together AI plans to expand its open-source AI inference capabilities for enterprise customers.
- The dedicated infrastructure is expected to become available in the first quarter of 2027.
Together AI, a San Francisco AI startup, has entered a $240 million multi-year agreement with IBM covering the deployment of NVIDIA AI infrastructure on IBM Cloud. The agreement gives Together AI access to dedicated computing infrastructure intended for large-scale AI inference.
The deal centers on the infrastructure required to run AI models after they have been developed. Inference refers to the process through which an AI system uses a trained model to generate outputs from new inputs. The new infrastructure is intended to support this activity at scale.
IBM Cloud will host the NVIDIA infrastructure covered by the agreement. Together AI will use the deployment to expand its ability to provide open-source AI inference services to enterprise customers.
The agreement therefore involves three major technology components: Together AI’s AI services, IBM Cloud’s cloud infrastructure and NVIDIA’s computing hardware. Together, those components will support the planned deployment for large-scale inference workloads.
The dedicated infrastructure is expected to become available in the first quarter of 2027. Until that deployment is available, the agreement represents a planned expansion of computing capacity rather than an already operational infrastructure installation.
For the San Francisco technology sector, the agreement places a locally based AI company at the center of a major infrastructure deployment involving two established technology providers.
San Francisco AI Startup Expands Inference Capacity
Together AI’s role in the agreement is focused on open-source AI inference. The company plans to use the additional infrastructure to provide AI services to enterprise customers that use open-source models.
Open-source AI models can be made available for broader use than proprietary systems controlled exclusively by one provider. Together AI’s planned deployment is specifically connected to inference, allowing those models to be used for AI workloads requiring computing resources after model development.
The $240 million agreement provides a defined financial commitment for the infrastructure arrangement. Its multi-year structure also establishes a longer-term relationship between Together AI and IBM rather than a short-term infrastructure purchase.
The agreement does not represent a change in Together AI’s San Francisco location. Instead, it expands the infrastructure supporting the company’s AI services through IBM Cloud.
Enterprise customers are a central part of the deployment. Together AI plans to use the infrastructure to expand services for organizations that require open-source AI inference capacity.
The focus on inference also distinguishes the agreement from an announcement centered on AI model development alone. The planned infrastructure is intended for the computing demands associated with operating AI models and generating outputs for users and applications.
The region’s startup ecosystem has also hosted events focused on AI infrastructure, applied artificial intelligence and enterprise deployment, including Silicon Valley AI startup discussions at TiEcon 2026.
NVIDIA Infrastructure Powers New AI Deployment
The planned deployment will use NVIDIA HGX B300 infrastructure. NVIDIA’s hardware will provide the computing foundation for the AI infrastructure being deployed through IBM Cloud.
The use of dedicated NVIDIA infrastructure gives the agreement a specific hardware component rather than leaving the computing requirements undefined. Together AI will use that infrastructure for large-scale AI inference under the multi-year arrangement with IBM.
NVIDIA infrastructure is therefore part of the operational structure of the agreement, while IBM Cloud provides the cloud environment and Together AI supplies the AI inference services.
The deployment is planned rather than immediately available. The dedicated infrastructure is expected to become available during the first quarter of 2027, establishing a future delivery point for the capacity covered by the agreement.
The infrastructure is intended to support open-source AI models and enterprise AI services. That purpose defines the primary use of the computing capacity under the agreement.
For technology companies, inference infrastructure is an important component of delivering AI systems to users. Once an AI model is available for use, computing resources are required to process requests and produce outputs. The Together AI agreement is specifically structured around that stage of AI deployment.
The arrangement also places hardware, cloud computing and AI services within one deployment. NVIDIA provides the infrastructure, IBM provides the cloud environment and Together AI uses the resulting capacity for its open-source AI inference services.
The hardware component is particularly relevant to the Bay Area’s AI ecosystem, where local coverage has also examined AI chiplet security concerns and the technical infrastructure supporting newer AI systems.
IBM Cloud Supports Open-Source AI Inference

Photo Credit: Unsplash.com
IBM Cloud is the infrastructure environment identified in the agreement. Together AI will deploy the NVIDIA HGX B300-based infrastructure through IBM Cloud to support its inference services.
The arrangement gives the San Francisco AI startup a dedicated infrastructure path for its planned expansion. The infrastructure is intended to serve large-scale workloads rather than a limited individual deployment.
Together AI’s use of IBM Cloud also places its open-source AI services within an enterprise-oriented cloud environment. The agreement specifically identifies enterprise customers as the audience for the expanded inference capabilities.
The $240 million value applies to the multi-year agreement between IBM and Together AI. The financial commitment covers the planned infrastructure arrangement and establishes the scale of the deployment.
The agreement’s structure provides a clear division of responsibilities. IBM supplies the cloud environment, NVIDIA supplies the computing infrastructure and Together AI uses that infrastructure to expand AI inference services.
The planned availability date in the first quarter of 2027 provides the next major milestone for the deployment. Until then, the infrastructure remains part of the announced expansion plan.
The agreement does not establish a new AI model from Together AI or IBM. Instead, it concerns infrastructure intended to support the operation of open-source AI models and the delivery of AI inference services to enterprise customers.
Other Bay Area technology developments have also involved AI applications beyond cloud infrastructure, including AI tools for San Francisco Bay, demonstrating the range of local projects using artificial intelligence for specialized applications.
Frequently Asked Questions
What is Together AI?
Together AI is a San Francisco-based AI startup focused on providing open-source AI models and AI inference services to enterprise customers.
What is the Together AI and IBM agreement?
Together AI and IBM signed a $240 million multi-year agreement to deploy NVIDIA AI infrastructure on IBM Cloud for large-scale AI inference.
How much is the Together AI IBM deal worth?
The multi-year agreement is valued at $240 million.
What NVIDIA infrastructure will Together AI use?
The planned deployment will use NVIDIA HGX B300 infrastructure hosted through IBM Cloud.
When is the new Together AI infrastructure expected to become available?
The dedicated infrastructure is expected to become available in the first quarter of 2027.







