Environment Planning
Environment planning establishes the infrastructure scope, ZCF management-plane topology, and node, network, and data-capacity requirements before deployment. The resulting plan provides the basis for executable node, network, capacity, and acceptance designs.
ZStack Cloud is the mandatory foundation of the current ZCF architecture. ZStack ZStone, ZStack Zaku, and ZStack ZNS provide storage, container, and network capabilities respectively; they can be deployed and connected to ZCF as needed by the construction stage and workload requirements. Installing ZCF does not replace deployment of the infrastructure components described above. This topic describes pre-deployment planning; for installation procedures, see Installation and Deployment.
Planning Overview
This topic defines the scope of environment planning, the information required in an implementation plan, and the usage boundaries of planning guidance. It establishes a consistent basis for environment planning.
Define the Environment Scope
A ZCF environment consists of deployed infrastructure and a new unified management plane. The management plane includes Unified Portal, Unified Identity and Access Management, ZCF Cloud Federation, the Observability and Operations Component, and installation and lifecycle management capabilities.
Existing ZStack Cloud environments can connect to ZCF without moving existing VMs or their associated resources, and can progressively gain a unified entry point, single sign-on, a unified asset view, and observability and operations capabilities. Version compatibility, network connectivity, available resources, and the change window must still be verified before onboarding.
Clarify Terminology Boundaries
To avoid ambiguity, this chapter uses the following terminology:
| Term | Meaning in This Chapter |
|---|---|
| Infrastructure environment | The runtime environment composed of deployed infrastructure components and their nodes, network, storage, and shared infrastructure services. |
| Infrastructure component | A component that is deployed, connected, and maintained according to its own version, such as ZStack Cloud, ZStack ZStone, ZStack Zaku, and ZStack ZNS. |
| Resource | A compute, storage, network, or data-capacity object provided or managed by an infrastructure component. |
Create an Implementation Plan
An environment plan should at least define:
- The deployment topology selection, either single-node or three-node high availability, and the acceptable scope of a management-service interruption.
- The placement of ZCF service nodes and the Observability and Operations Component, their resource budget, and resource headroom after a fault.
- Access relationships for node addresses, the ZCF service VIP (stable service endpoint), the authentication endpoint, and infrastructure component connection endpoints.
- Metric and log collection scope, daily growth, retention, and storage capacity.
- External services, change windows, and owners required for installation, incident handling, and upgrade.
Planning Guidance and Usage Boundaries
In this topic, prerequisites are conditions required before deployment, recommendations are production-oriented design choices, and field estimates are parameters that must be determined from actual workloads. Reference configurations and capacity calculations explain planning methods; they are not performance guarantees or capacity limits.
Deployment Topology and Methods
This topic defines the node roles, deployment topology, Observability and Operations placement, and deployment entry conditions for ZCF. It provides the basis for subsequent resource, network, and implementation design.
Plan Node Roles
| Role | Primary Responsibility | Planning Focus |
|---|---|---|
| Installer node | Runs Installer and provides the installation wizard. | Read access to the local package directory, SSH connectivity to ZCF service nodes, and installation permissions. |
| ZCF service node | Runs services for the ZCF management plane. | Node count, CPU architecture, node IP address or FQDN, SSH username and port, ZCF service VIP, failure domains, and operations. |
| Observability and Operations node | Hosts observability and operations capabilities and related data processing. | By default, it is co-located with ZCF service nodes; a separate node may be configured according to the installation plan. Compute and data storage are sized separately. |
| ZStack Cloud management node | Provides management connectivity for a deployed ZStack Cloud environment. | Platform reachability, management-node mode, IP address or FQDN, access point, and connection information. |
| Other infrastructure components | Provide storage, container, or network services. | After deploying the corresponding component, confirm the connection address, version, and collection conditions. |
Node roles define deployment responsibilities and do not imply a requirement for new physical servers. The deployment relationship between the Installer node and ZCF service nodes, and whether Observability and Operations uses separate nodes, must follow the installation plan for the target version.
Select a Deployment Topology
- Single-node deployment:
- Concentrates ZCF services on one node and is suitable for evaluation, testing, demonstration, and noncritical environments where management-service interruptions are within the acceptable range.
- Metric ingestion, log retrieval, and management requests may share compute and disk resources.
- Recovery arrangements and maintenance windows should be defined in advance for node maintenance or failure.
- Three-node HA deployment:
- Uses three ZCF service nodes to form a high-availability unit and provides a ZCF management-service VIP; it is suitable for production environments.
- Service nodes should be distributed across physical failure domains, and address conflicts, network reachability, and failover conditions should be verified. Resource headroom should be reserved for recovery and maintenance.
- High availability for the ZCF management plane should be planned separately from high availability for ZStack Cloud and other infrastructure components.
Determine Observability and Operations Placement
ZCF Observability and Operations is provided by ZCF service nodes by default and can be co-located with them. When metric and log collection, historical queries, or reporting workloads increase, or when resource isolation, data protection, and availability requirements need to be addressed separately, separate nodes can be planned according to the installation plan for the target version.
- Co-locate with ZCF service nodes:
- The initial data volume is small, and collection, query, and management workloads fit within the same resource budget.
- Fewer nodes; monitor contention from logs and metrics against management services.
- Use separate nodes:
- Data ingestion, historical queries, or reporting workloads are high and need independent resource planning.
- Adds nodes, connections, and maintenance work; data protection and availability must be confirmed for the target version.
Select a Deployment Method
ZCF provides two deployment entry points: Installer and the ZStack Cloud Application Market. The entry point can be selected according to the existing infrastructure environment and target deployment topology.
- Installer: Suitable when target service nodes are prepared and a wizard will configure node and deployment relationships.
- ZStack Cloud Application Market: Suitable when an available ZStack Cloud environment will host ZCF as VMs.
Both entry points require architecture-matched media, nodes, and ZStack Cloud management connection information in advance; parameters and procedures must be confirmed in the target-version installation documentation.
Plan an Implementation Path
The implementation path should be determined by the build status of the infrastructure environment.
- New environment: Build ZStack Cloud → Prepare other infrastructure components as needed → Deploy ZCF → Configure component connections and authentication → Verify asset synchronization and observability data.
- Existing ZStack Cloud environment: Confirm that ZStack Cloud is available and prepare deployment resources → Install ZCF from the Application Market → Complete component onboarding and unified authentication configuration → Verify asset synchronization and observability data.
Node and Capacity Planning
This topic assesses the node and data resources required by ZCF based on onboarding scale and Observability and Operations workloads. It produces an actionable capacity plan.
Collect Planning Inputs
ZCF management load depends on the scale of onboarded infrastructure components and the number of resources they manage. It is also affected by collection granularity, log rate, historical-query scope, and the number of users. Planning should collect at least the following information:
- Infrastructure components and versions: component combinations, target versions, and onboarding scope.
- Onboarding scale: the number of onboarded infrastructure components, VMs, and other managed resources, as well as simultaneous active users.
- Observability and Operations workload: metric scope and sampling intervals, average and peak log growth, historical-query scope, concurrent-query patterns, and scheduled reports.
- Service objectives and growth expectations: data-retention requirements, availability requirements, and expected growth during the planning period.
Define an Initial Scale
Initial planning uses VM count and simultaneous active users as scale bands:
| Scale | VM Count | Simultaneous Active Users |
|---|---|---|
| Small | Up to 200 | Up to 5 |
| Medium | 201 to 500 | Up to 10 |
| Large | 501 to 1,000 | Up to 20 |
Note: Scale bands select validation workloads; they do not replace node capacity sizing or constitute capacity limits. When VM count and user concurrency indicate different bands, use the higher band for initial assessment. Environments beyond these bands, or with high log volume, intensive historical queries, or long retention periods, require a dedicated assessment. Actual resources, metric series, and data volumes from onboarded ZStack Zaku, ZStack ZStone, and ZStack ZNS environments must be counted separately.Create a Node Specification
The final configuration should be recorded per node and include at least the following fields:
| Field Category | Information to Record |
|---|---|
| Node and role | Node ID, hosted components, quantity, and applicable workload. |
| Compute and storage | CPU architecture, CPU or vCPU count, memory, system disks, data disks, and media and I/O requirements. |
| Network and operating system | NICs and bandwidth, management address, and operating-system version. |
| Availability and data | Host or physical failure domain and data-retention conditions. |
Note: Shared nodes are counted once as physical or virtual resources, with component budgets aggregated within the node. Instance specifications in Application Market installation documentation may serve as an initial resource reference, but they must not be used to infer production capacity, management-resource limits, or separate Observability and Operations node specifications. Node configurations for ZStack Cloud, ZStack ZStone, ZStack Zaku, and ZStack ZNS follow the deployment specifications for their respective components and versions.Validate Production Sizing
Production sizing should be determined through representative workload validation, including:
- Validation scenarios: normal collection, log peaks, concurrent user queries, and scheduled report execution.
- Recorded indicators: CPU, memory, disk throughput, and response behavior.
- High-availability validation: workload distribution after node failure in an agreed environment.
When historical queries and log ingestion are the primary source of pressure, increasing data-processing and storage resources is often more effective than simply increasing management-node count. The final design should identify the role, resource budget, and expansion trigger for every node.
Network and Access Planning
This topic defines the network conditions required for ZCF deployment, management, infrastructure component onboarding, observability data transfer, and shared infrastructure services. It produces verifiable access paths, service endpoints, and bandwidth planning.
Map Access Relationships

The network design must cover administrator access, installation and deployment, node-to-node communication, infrastructure component onboarding, observability data, and shared infrastructure services. Small environments may reuse a physical network. Production environments should separate paths and budget bandwidth according to access targets and traffic characteristics, so bulk log transfer does not contend with critical management traffic.
Plan Addresses and Service Endpoints
Record and verify the following address and service-endpoint information separately:
- Node management addresses and service endpoint: Every ZCF service node requires a stable management address. Node IP addresses are used for node connection and maintenance, while the service VIP is the stable access endpoint; record them separately.
- DNS names and routing: When using FQDNs, verify name resolution and routing from administrator endpoints, the Installer, and service nodes.
- High-availability VIP: Before implementation, a high-availability deployment must confirm the VIP subnet, hosting interface, reserved address, and network-device support for failover.
- Authentication and component entry points: ZIAM authentication entry points and infrastructure component access addresses should remain consistent and stable. Unified identity authentication does not automatically map roles or permission models between components.
Create an Access Matrix
| Source → Target | Purpose | Planning Requirement |
|---|---|---|
| Administrator endpoint → Installer node | Installation wizard | Confirm the protocol and port in the target-version startup configuration. |
| Installer node → ZCF service nodes / ZStack Cloud management nodes | Deployment, precheck, and installation integration | Confirm SSH or the actual management port, source subnet, and permissions. |
| Administrator endpoint → ZCF management entry point, identity authentication entry point, and infrastructure component entry points | Management access, login, and redirects | Confirm the protocol, port, DNS name, certificate, and browser reachability for each access path. |
| ZCF service nodes → onboarded infrastructure component management endpoints | Connection validation, resource synchronization, and collection configuration | Confirm the access required by each component's actual interfaces. |
| Infrastructure component collection endpoints → ZCF Observability and Operations nodes | Metric and log transfer | Confirm the push or pull direction, endpoint, and peak bandwidth for each component. |
| Installer nodes / ZCF service nodes → shared infrastructure services | Name resolution, time synchronization, and media access | Confirm reachability and ownership of DNS, time services, certificates, and media sources. |
This matrix provides an access-planning framework and does not constitute a complete firewall allowlist. The final allowlist must identify component ports, protocols, directions, and source subnets, and must follow the target-version installation documentation.
Size Bandwidth and Shared Services
- Observability data bandwidth: Estimate bandwidth from peak transfer volume and include protocol overhead, potential retransmission, and concurrent queries. A reachable management page demonstrates only that one path is available; it is insufficient to establish that all deployment and collection paths meet requirements.
- Time, certificates, and proxies: All management and collection nodes should use a consistent, reliable time source. When using DNS names and encrypted access, verify certificate name matching, trust chains, validity periods, and renewal ownership. If an external proxy is deployed, confirm that it is a supported access method for the target version.
Observability Data and Storage Planning
This topic estimates the storage requirements for metrics, logs, and runtime data related to Observability and Operations. It produces a plan for data capacity, storage resources, and ongoing review.
Size System and Data Space Separately
- Disk budget scope: The ZCF disk budget includes the operating system, component runtime data, installation and upgrade media, metrics, logs, and operating and maintenance headroom.
- Co-location impact: When Observability and Operations capabilities share nodes with other management components, pay particular attention to contention for the same disks.
- Storage-performance validation: Sufficient capacity does not demonstrate that write latency and query throughput meet requirements; storage media and I/O capability must also be validated against actual workloads.
Estimate Metric and Log Capacity
- Metric capacity: It depends mainly on active time-series count, sampling interval, retention period, and actual storage overhead. Daily samples can be estimated as
active time series × 86,400 ÷ sampling interval in seconds, then combined with actual per-sample storage, indexes, and runtime overhead from a representative environment. - Log capacity: Prefer measured storage growth after the actual collection scope is enabled. Log types and collection coverage must be checked for each infrastructure component; do not assume that all system and application logs enter the log center automatically after a component is onboarded.
- Steady-state data space: For data with a defined effective retention period, estimate it as
daily metric storage growth × metric retention days + daily log storage growth × log retention days. Record whether measurements include indexes and replicas. Net disk growth after reaching steady state must not be treated directly as daily write volume.
Convert to Physical Capacity and Revalidate
- Physical-capacity conversion: Add system space, temporary upgrade space, data-maintenance space, and safety headroom to the calculated result, then convert it to node and underlying-storage capacity based on the actual storage layout. A three-node high-availability deployment does not imply that all historical data automatically has three replicas, nor that all disk capacity is available for unique data.
- Ongoing review and expansion: After go-live, observe data growth, free disk space, write latency, and query behavior continuously. Measure daily growth again and update forecasts when infrastructure components are added, collection scope changes, or sampling frequency increases. Schedule expansion before capacity is exhausted. Backup and online retention serve different purposes and cannot replace each other.
Pre-deployment Checks and Acceptance Preparation
This topic defines pre-deployment environment boundaries, checks, and acceptance scope. It produces an actionable precheck and handover preparation checklist.
Define Environment Boundaries and Operations Ownership
- Environment isolation: Production, validation, and demonstration environments should manage access endpoints, authentication configurations, infrastructure component connections, and data-retention requirements separately. The validation environment should reproduce key infrastructure component versions and connection methods for installation, upgrade, authentication, and collection integration, rather than applying experimental operations directly to production.
- Planning scope and lifecycle: This topic uses a single ZCF environment and its supported infrastructure component onboarding scope as its planning boundary. Multiple environments may have separate node and network designs. Cross-site unified onboarding, cross-site high availability, and disaster-recovery failover require separate design after supported boundaries are confirmed. ZCF manages the lifecycle of its management components; ZStack Cloud, ZStack ZStone, ZStack Zaku, and ZStack ZNS continue to follow their own supported maintenance and upgrade procedures.
Complete Pre-deployment Checks
| Check | Completion Criteria |
|---|---|
| Infrastructure components and versions | The component combination, ZCF media, and target-version compatibility conditions are defined. |
| Infrastructure environment | ZStack Cloud is available and infrastructure components to be onboarded are prepared. |
| Deployment method and topology | Installer or Application Market, and single-node or three-node high availability, are defined. |
| Nodes and network | Compute, storage, and fault headroom are sized; node addresses, VIP, DNS names, and required paths are complete and reachable. |
| Infrastructure services and permissions | Name resolution, time synchronization, certificate conditions, and deployment accounts are verified. |
| Data and operations | Collection scope, effective retention, storage budget, maintenance window, backup and recovery, and owners are defined. |
Installer prechecks help identify connectivity, resource, and configuration issues. Production capacity, failure domains, and recovery objectives must still be validated against the environment plan.
Prepare Layered Acceptance
- Acceptance scope: Acceptance should cover installation results, management entry points, identity authentication, infrastructure component connections, observability data, alerts and views, availability, and operations handover.
- Acceptance basis: Historical records in the asset view do not prove that current collection is normal. Validate each onboarded component and each enabled data scope.
The final handover should include node, network, and capacity specifications, version records, and a recovery plan. With these preparations in place, delivery teams use prechecks and layered acceptance to complete handover preparation for a maintainable ZCF environment.
