Begin with the workload, not a GPU count
An enterprise AI plan should first establish how the model will be used. Training, fine-tuning an existing model and serving production inference have different capacity and operational requirements. Concurrent requests, model memory, context length and acceptable response time need to be evaluated together. A catalog description of powerful AI infrastructure cannot replace these decisions. Business owners should explain which activity needs improvement, while platform teams identify how that improvement will be measured. Cisco's separate training and inference design guides support this distinction: a reference architecture should not be applied unchanged to every workload simply because each one uses GPUs.
Separate frontend and backend responsibilities
Cisco AI POD references use the frontend network for application access, storage, management and associated control traffic. The backend network addresses communication between nodes in distributed GPU workloads. Separate fabrics make different traffic requirements visible, but do not mean that every inference application needs the same backend design. If a model fits on one GPU or uses a different serving arrangement, revisit the corresponding architecture decision. The diagram separates spine and leaf groups for the two fabrics. Network speeds shown in the source are connection labels for that example; they are not measured performance results for an organization's application.
Cisco AI POD: frontend and backend fabrics
Select nodes to read their roles; zoom or enter full screen to follow the connections.
- Backend fabric (E-W)
- Frontend / storage (N-S)
- ND bağlantıları / ND connections
Select a node to read its role. Drag to pan or use the buttons to zoom.
The August 2026 inference reference keeps separate backend and frontend fabrics. First and N compute nodes, management cluster and storage modules follow the illustrated scope; intermediate nodes remain an ellipsis. This is not a universal hardware recipe for all AI POD variants. Backend connectivity is represented at compute-group scope; individual NIC/port assignments are not inferred from the first/N icons.
Match compute and GPU choices to model behavior
Cisco UCS and NVIDIA GPUs provide the compute foundation of the selected reference architecture. Platform selection should consider model size, precision or quantization, parallelism requirements and the expected usage profile. Do not assume that a benchmark obtained with one model and dataset will transfer unchanged to another. During a pilot, record latency distributions and behavior at capacity as well as average throughput. The first and N compute nodes in the drawing indicate that additional members can exist between them; they do not prescribe a purchasing quantity. Choose hardware by considering supported components alongside measurements from the intended workload.
Treat storage as part of the model lifecycle
Storage on an AI platform is more than a location for model files. Datasets, model versions, checkpoints and enterprise context used by the application can have different access requirements. In a RAG application, document freshness and permissions also influence the result. Identify data ownership, access authorization, retention needs and the model-update process early in design. The FlashStack source topology illustrates specific storage components; an alternative should not be assumed to have identical performance or support characteristics. Acceptance should separately observe model loading, access to required data and controlled transition to an updated version, with responsibility assigned to each part of the lifecycle.
Define the software and serving arrangement
The software that serves the model matters as much as available GPU capacity. Cisco's inference reference treats OpenShift and model-serving components as part of the platform design, with NVIDIA AI Enterprise providing relevant software and support options. An organization's Kubernetes, OpenShift or MLOps choice should be evaluated together with version compatibility, operating responsibility and the update plan. Exposing a model through an API also requires decisions about application identity, quotas and access. The reference drawing does not show every software API path, so those paths are not invented as physical cables. Application and infrastructure teams need a shared, concrete serving scenario.
Cisco AI Networking: workload-appropriate Ethernet fabric
A Cisco AI Networking decision should follow how the GPU workload is distributed. Communication patterns for training and distributed inference can differ from frontend access. Evaluate a spine–leaf or rail-optimized option against the chosen servers, NICs, GPUs and software arrangement. The catalog's performance and latency goals are requirements to measure with the actual workload in a pilot. Cisco's training guide describes different backend options, while the FlashStack inference topology shown here is a specific reference. Do not interchange hardware and connection labels between those sources: associate them with support conditions for the selected platform.
Connect Secure AI Factory to specific protection needs
Cisco Secure AI Factory with NVIDIA brings compute, networking, security and relevant AI software into a combined approach. An AI POD infrastructure decision is nevertheless distinct from securing the AI application itself. Network segmentation, workload protection and inspection of model inputs address different scopes. Establishing the risks to be addressed by AI Defense or Hypershield is more useful than placing every product on the same diagram. Hardware, software, licensing and deployment choices determine the actual integration scope. Security requirements should therefore be designed with application and data flows, rather than appended as a note after the capacity plan has already been agreed.
Assign operational ownership across management tools
Server, network and application teams need different views of the same platform. Intersight focuses on UCS management, while Nexus Dashboard addresses network-fabric management and operations. Cisco Cloud Control also appears in the source topology; it is not an additional data-plane hop for application packets. The operating model should identify who evaluates each alarm, which measurements are considered together when capacity is constrained, and who approves version changes. A successful platform is more than a running cluster. Adding models, updating them, rolling back an unsuitable version and expanding capacity all require an understandable distribution of responsibility between the participating teams.
Turn pilot evidence into a production decision
A proof of concept should produce workload measurements and understood limits, rather than just a successful demonstration response. Recording the dataset, model version, concurrency and measurement duration makes results comparable later. Add the organization's important scenarios—such as a network or node failure, model update and service restart—to acceptance planning. Trustnet's catalog describes analysis, design, implementation and support as the delivery approach for these decisions. A production decision becomes concrete when an appropriate reference architecture, pilot evidence and operational ownership are ready together. The diagram helps explain those decisions; it does not independently establish a capacity figure or continuity commitment.



