Skip to main content
Health checks detect hardware and software issues on your GPU nodes before they impact running workloads. Together supports two types of health checks: active checks that run on-demand diagnostic tests, and passive checks that monitor nodes continuously in the background. Issues detected by either type can trigger node repair recommendations.

Active health checks

Active health checks let you run targeted diagnostic tests on your GPU nodes to validate GPUs, InfiniBand networking, storage, and other components. These tests require full GPU utilization and run on demand.

How to run health checks

Quick steps

  1. Navigate to your cluster in the Together web interface.
  2. Go to the Cluster Details tab and select the Health Checks sub-tab.
  3. Select the Run a health check button (top right).
  4. In the Run Health Checks dialog, select one or more tests to run:
    • DCGM Diag: NVIDIA GPU diagnostics.
    • GPU Burn: GPU stress test.
    • Single-Node NCCL: Single-node GPU communication test.
    • NVBandwidth: CPU to GPU Bandwidth: PCIe bandwidth test.
    • NVBandwidth: GPU to CPU Bandwidth: PCIe bandwidth test.
    • NVBandwidth: GPU-CPU Latency: PCIe latency test.
    • InfiniBand Write Bandwidth: InfiniBand network performance test.
    • Storage Performance (Preview): Storage throughput and data-integrity test.
  5. Select Next: Select Nodes.
  6. Choose which nodes to test.
  7. (Optional) Configure test parameters like duration or diagnostic level.
  8. Select Run to start the health checks.
Active tests: These health checks require full GPU utilization from the node and impact any running workloads during the test.

Available tests

Each health check validates different aspects of your GPU infrastructure:

GPU diagnostics

DCGM Diag
  • Runs NVIDIA Data Center GPU Manager diagnostics.
  • Validates GPU compute capability, memory integrity, and thermal performance.
  • Configurable: Diagnostic level (1-3, where 3 is most comprehensive).
  • Use for: Comprehensive GPU health validation.
  • Learn more: NVIDIA DCGM documentation.
GPU Burn
  • Stress tests GPUs with intensive compute workloads.
  • Validates stability under sustained high utilization.
  • Configurable: Test duration.
  • Use for: Identifying thermal issues, power problems, or instability.
  • Learn more: GPU Burn on GitHub.

Network performance

Single-Node NCCL
  • Tests NVIDIA Collective Communications Library on a single node.
  • Validates GPU-to-GPU communication within the node.
  • Use for: Multi-GPU training readiness.
  • Learn more: NVIDIA NCCL documentation.
InfiniBand Write Bandwidth
  • Measures InfiniBand network write throughput.
  • Validates high-speed interconnect performance.
  • Use for: Distributed training and multi-node workloads.

PCIe performance

NVBandwidth Tests
  • CPU to GPU Bandwidth: Host-to-device transfer rates.
  • GPU to CPU Bandwidth: Device-to-host transfer rates.
  • GPU-CPU Latency: Data transfer latency.
  • Use for: Identifying PCIe bottlenecks or degraded lanes.
  • Learn more: NVIDIA nvbandwidth documentation.

Storage

Storage Performance
  • Runs storage benchmarks against shared and local storage tiers available to the cluster.
  • Validates data integrity with a checksummed write and read-back, and measures sequential read and write throughput.
  • Use for: Detecting slow or degraded storage before it bottlenecks data loading or checkpointing.
Preview: The Storage Performance check is in preview. Reach out to support for access.

Training performance

TorchTitan Training
  • Runs a short TorchTitan training benchmark on a single node or across multiple nodes.
  • Measures the steady-state median model FLOPs utilization (MFU), skipping the warmup step.
  • Validates that a node sustains the expected training throughput for its GPU type, catching grossly degraded nodes that still pass point-in-time diagnostics.
  • Requires: NVIDIA B200 (Blackwell) GPU nodes. The benchmark runs a CUDA and PyTorch training stack (PyTorch with FlashAttention 4) built for the B200 architecture (compute capability sm_100), so it does not run on other GPU types.
  • Runtime: About 15 minutes total, roughly 9 minutes to initialize and 6 minutes to run.
  • Use for: End-to-end training-readiness validation.
Preview: The TorchTitan training check is in preview. It currently runs on demand on B200 nodes only, is not part of automatic acceptance testing, and its behavior and thresholds may change.

Understanding test results

Health check results are displayed in the Health Checks table:
  • Status: Passed (green) or Failed (red) indicator.
  • Last Run: Timestamp of test execution.
  • Node Tested: Which nodes were included in the test.
  • Details: Select View details to see:
    • Full test output.
    • Detailed metrics and measurements.
    • Workflow CR (Custom Resource) with complete results.
    • Pass/fail criteria details.

Pass/fail thresholds

Performance-based health checks compare the measured result against a reference value. For bandwidth tests, a test passes when the measured value is greater than or equal to the threshold. For latency tests, lower is better, so a test passes when the measured value is less than or equal to the threshold. The following thresholds are the defaults applied during health checks and automatic acceptance testing.
Units differ by test: NCCL and NVBandwidth bandwidth thresholds are reported in gigabytes per second (GB/s), while InfiniBand write bandwidth is reported in gigabits per second (Gb/s). The two are not directly comparable (320 Gb/s is 40 GB/s). Latency is reported in nanoseconds (ns), and TorchTitan MFU is reported as a percentage.
The TorchTitan pass threshold and reference MFU depend on the model and node topology. The 30% default is a loose floor for the Llama 3 8B single-node case (reference MFU is about 36.5%). Larger models and multi-node runs use different reference values, which are set per run. For example, Llama 3 70B reaches about 38% MFU on a 4-node cluster. These thresholds are tuned for current GPU node types and may be adjusted over time as hardware and reference baselines change.

Automatic acceptance testing

When you provision a new GPU cluster, Together automatically runs acceptance tests on each node before making it available for your workloads. This ensures that all nodes meet quality standards before joining your cluster.

During cluster provisioning

The cluster provisioning process includes an automatic testing phase: Phase: Running Tests During this phase, each node undergoes single-node acceptance tests:
  • DCGM Diag Level 2: Comprehensive GPU diagnostics.
  • 5-minute GPU Burn: Sustained GPU stress test.
  • Single-Node NCCL: GPU-to-GPU communication validation.
  • Multi-Node NCCL: GPU-to-GPU communication validation across node GPUs.
  • Storage Performance: Storage performance validation for storage volumes attached to the cluster.
You’ll see the cluster status as:
  • Running Tests: Acceptance tests are in progress.
  • Tests Failed: One or more acceptance tests did not pass.
  • Running: Tests passed and the cluster is ready.

Viewing acceptance test results

If acceptance tests fail during provisioning:
  1. Navigate to your cluster in the Together Cloud UI.
  2. Go to the Cluster Details tab.
  3. Select the Health Checks sub-tab.
  4. Find the acceptance test runs for the affected nodes.
  5. Select View details to see:
    • Which specific test failed (DCGM Diag, GPU Burn, NCCL, or Storage Performance).
    • Detailed error messages and logs.
    • Performance metrics from the tests.
Automatic remediation: If acceptance tests fail, Together’s infrastructure team is automatically notified and investigates. Nodes that fail acceptance tests are not added to your cluster until the issue is resolved.

Why acceptance testing matters

Automatic acceptance testing provides several benefits:
  • Quality assurance: Every node is validated before you can use it.
  • Early detection: Hardware or configuration issues are caught immediately.
  • Reduced downtime: Problems are fixed before they impact your workloads.
  • Consistent performance: All nodes meet the same performance standards.
Provisioning time: Acceptance tests typically add 5-10 minutes to cluster provisioning time, but this ensures you receive fully validated, production-ready nodes.

When to run active health checks

Proactive testing:
  • Before deploying critical workloads.
  • After cluster scaling events.
  • On a regular schedule (weekly or monthly).
  • After maintenance windows.
Reactive testing:
  • When experiencing unexplained job failures.
  • Before triggering node repair actions.
  • When investigating performance degradation.
  • After node repairs to validate fixes.
Specific issue investigation:
  • Training instability: Run GPU Burn and DCGM Diag.
  • Slow data loading: Run Storage Performance and NVBandwidth tests.
  • Multi-GPU failures: Run Single-Node NCCL.
  • Distributed training issues: Run InfiniBand tests.
  • Low training throughput: Run TorchTitan Training.

Best practices

  1. Schedule workload-free windows: Health checks require full GPU utilization.
  2. Start with DCGM Diag: It provides a comprehensive overview of GPU health.
  3. Run baseline tests: Test new nodes immediately to establish a performance baseline.
  4. Document results: Keep records of passed tests for comparison.
  5. Test after repair: Always validate node health after repair actions.
  6. Use appropriate test levels: Higher DCGM diagnostic levels take longer but are more thorough.
Workload impact: Health checks fully utilize the GPU and can interfere with running workloads. Run tests during maintenance windows or on idle nodes.

Passive health checks

Passive health checks run continuously on every node in your cluster, observing real workloads, system logs, and GPU metrics in the background. Unlike active checks, passive checks require no manual action and have zero impact on running jobs.

How passive checks work

Passive checks monitor node telemetry and system logs in real time. There is no synthetic load: the system observes your production traffic and flags degradation as it happens. When an issue is detected, the system creates an internal alert with supporting evidence. If the alert meets repair criteria, a repair recommendation is generated automatically.

Detected failure modes

Passive checks continuously watch node telemetry, GPU metrics, and system logs, and raise a signal when a node shows one of the conditions below. Each signal that meets repair criteria generates a repair recommendation with a suggested action.

GPU and accelerator

InfiniBand networking

Node OS, kernel, and resources

Scheduler

See Recommended repair actions for the repair action each signal maps to.
Detection coverage is expanding, and automated recommendations are enabled per cluster. Not every signal triggers an automated recommendation today; some raise an internal alert that Together’s team reviews. Future releases will cover additional hardware and software signals.

Active vs. passive checks

  • Active checks: Run on demand, use synthetic workloads that require full GPU utilization, and validate specific hardware capabilities. Best for targeted diagnostics and pre-deployment validation.
  • Passive checks: Run continuously with zero overhead, observe real workloads and system logs, and catch degradation as it happens. Best for ongoing monitoring and early detection.
Both types feed into node repair recommendations.

Next steps

Node repair

Restore unhealthy nodes through automated recommendations or manual repair actions.

Cluster management

Manage, monitor, and scale your GPU clusters.