Article Details

Huawei Cloud USDT Top-up Huawei Cloud ECS scheduled restart script

Huawei Cloud2026-05-15 14:32:36TrustCloud

Why You’d Even Want a Scheduled ECS Restart

Let’s be honest: nobody wakes up eager to restart their servers. Servers rarely send cheerful emails like, “Good morning! Today we will be gently rebooting ourselves for your convenience!” They just sit there, quietly doing important work until something changes—patches, kernel updates, networking weirdness, or the occasional cosmic ray event (just kidding… mostly).

A scheduled restart is one of those boring-but-effective operational practices. It can help with:

  • Applying certain updates that require a reboot to fully take effect.
  • Clearing lingering memory leaks or stuck processes—yes, even the ones you swear you cleaned up.
  • Resetting network state when you notice long-term weirdness that refuses to reproduce on demand.
  • Reducing “mystery time”: if you know your system restarts on schedule, you can correlate behavior with reboot events instead of guessing.

However, a scheduled restart can also be a fast way to accidentally create an outage if you’re careless. The key is to script responsibly: target the right instances, respect time windows, log everything, and add safeguards so the script behaves like a responsible adult rather than a caffeinated raccoon.

What This Article Covers

You’re asking about a “Huawei Cloud ECS scheduled restart script.” So we’ll build a solution that generally works like this:

  • You define which ECS instances should be considered for restarts.
  • You schedule the script using cron or a similar scheduler.
  • The script checks whether it’s the right time to restart (and optionally whether it’s safe).
  • The script triggers a restart for the chosen instances through Huawei Cloud API calls or a command tool approach.
  • It records actions in logs, supports dry-run mode, and includes safety locks.

Because exact tooling varies (for example, whether you use Huawei Cloud CLI, REST API directly, or a specific SDK), this article focuses on the design principles and a practical command structure. You’ll adapt the “call” portion to match your environment. Think of it as the skeleton; you bring the meat (your credentials and exact API endpoint details).

High-Level Design: Don’t Write a Restart Script, Write a Restart Plan

Before code, define the rules. A good restart script answers these questions:

  • Which instances? By instance ID list, by tags, by naming convention, or by a “restart-group” label.
  • When? A fixed time window (like 03:00–04:00 local) or a “only if it hasn’t restarted in X days” rule.
  • How? Hard restart or soft reboot (if supported), and whether you pause workloads first.
  • What is “safe”? Is there a health endpoint check? Is there a maximum number of instances allowed to restart simultaneously?
  • What if something fails? Partial failure handling, retries, and alerting.
  • How will you know what happened? Logs with timestamps, instance IDs, and result codes.

If you skip these, your script becomes a “spray and pray” button. We prefer “controlled, repeatable, auditable.”

Common Mistakes (So You Can Avoid Them)

  • Restarting the wrong instance: You typed the wrong ID, or your filter is too broad. Fix with explicit allowlists and clear naming.
  • No dry run: Always test the selection logic in a dry run before actually restarting anything.
  • No rate limiting / concurrency control: Restarting ten instances at once can take out redundancy depending on your architecture. Add a “max concurrent restarts” setting.
  • Time zone confusion: Scheduled tasks use server time. Confirm whether cron is running in UTC or local time. Your future self will thank you.
  • Not logging: If you can’t prove what the script did, you’re not operating—you’re hoping.
  • Ignoring failures: If API calls fail, your script should log errors and stop or retry intelligently.

Prerequisites You’ll Need

At minimum, you need:

  • Access to Huawei Cloud credentials capable of calling ECS restart actions (via API or CLI).
  • An environment to run the script: a management host, CI runner, or even one of your ECS instances (though beware of circular dependencies).
  • A scheduler: cron (common), systemd timers, or cloud-native scheduling if you’re using it.
  • Networking reachability to Huawei Cloud API endpoints (if calling directly).

Huawei Cloud USDT Top-up Also recommended:

  • Logging directory and permissions.
  • A lock mechanism to prevent overlapping runs.
  • Optional health check hooks (for example, query an internal endpoint before restarting more nodes).

Decision: Restart by Instance ID, or by Tags?

There are two common approaches:

Approach A: Explicit allowlist of Instance IDs

You maintain a list of instance IDs (or a mapping by environment). This is the simplest and safest. The downside is manual updates when instances change.

Approach B: Select instances by tag or filter

You query instances that have a tag like restart-group=nightly. This scales better, but increases risk if tag logic is wrong. To keep it safe, you still want a dry-run mode and conservative filters.

For most teams, start with Approach A. Once it works reliably, graduate to tags if you need automation at scale.

Scheduling Strategy: Daily, Weekly, or “Restart If Needed”

Pick a cadence that matches your risk tolerance and system behavior:

  • Daily restart: Good for short-lived systems or where you actively want frequent refresh. Riskier if you don’t have redundancy.
  • Weekly restart: Often the sweet spot for maintenance windows.
  • Monthly restart: For stable systems where you primarily want to clear edge-case issues.

You can also add a rule: “Restart if last restart was more than X days ago.” This reduces unnecessary reboots. It’s like changing your car’s oil—eventually you do it, but you don’t want to do it just because the calendar looks moody.

Shell Script Template (Practical and Safe)

Below is a robust shell script template that demonstrates:

  • Huawei Cloud USDT Top-up Config variables at the top
  • Dry-run mode
  • Lock to prevent overlaps
  • Logging to file
  • Placeholder API/CLI restart call

Because Huawei Cloud authentication and API invocation differ depending on your setup, the “RESTART” function includes a placeholder. Replace that portion with the exact command or API call you use in your environment.

Example: Bash Scheduled Restart Script

Copy and adapt the snippet. It’s written to be readable and “future you” friendly.

Script: huawei_ecs_scheduled_restart.sh

#!/usr/bin/env bash
set -euo pipefail

# ============================
# Configuration
# ============================

# If DRY_RUN=1, the script will log what it WOULD do but will not restart.
DRY_RUN=1

# Max number of instances to restart in one run.
MAX_CONCURRENT_RESTARTS=2

# Lock file to prevent overlapping runs.
LOCK_FILE="/tmp/huawei_ecs_restart.lock"

# Log file
LOG_FILE="/var/log/huawei_ecs_restart.log"

# A simple allowlist of instance IDs.
# Replace with your real instance IDs.
INSTANCE_IDS=(
  "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
  "yyyyyyyy-yyyy-yyyy-yyyy-yyyyyyyyyyyy"
)

# Optional: environment/region identifiers for your API calls.
# REGION="cn-north-4"
# PROJECT_ID="..."

# Optional: a time window check (in local time).
# Example: Only restart if current hour is between 2 and 5 inclusive.
ALLOW_START_HOUR=2
ALLOW_END_HOUR=5

# ============================
# Helpers
# ============================

log() {
  # Logs with timestamp
  local msg="${1}"
  echo "$(date '+%Y-%m-%d %H:%M:%S') ${msg}" | tee -a "$LOG_FILE"
}

in_time_window() {
  local hour
  hour="$(date '+%H')"
  # Bash: hour is string. Convert to number.
  hour=$((10#$hour))

  if [[ $hour -ge $ALLOW_START_HOUR && $hour -le $ALLOW_END_HOUR ]]; then
    return 0
  fi
  return 1
}

acquire_lock() {
  # Use flock if available; fallback to mkdir trick.
  if command -v flock > /dev/null 2>&1; then
    exec 200>"$LOCK_FILE"
    flock -n 200
    return $?
  fi

  if mkdir "$LOCK_FILE" > /dev/null 2>&1; then
    return 0
  fi
  return 1
}

release_lock() {
  if command -v flock > /dev/null 2>&1; then
    # flock releases when file descriptor closes
    return 0
  fi
  rmdir "$LOCK_FILE" 2>/dev/null || true
}

restart_instance() {
  local instance_id="$1"

  if [[ "$DRY_RUN" == "1" ]]; then
    log "DRY_RUN=1: Would restart instance: ${instance_id}"
    return 0
  fi

  # ==========================================================
  # TODO: Replace this placeholder with your real restart call.
  # ==========================================================
  # Example placeholder options:
  # - Use Huawei Cloud CLI (if installed and configured)
  # - Call Huawei Cloud ECS API with curl + OAuth/token
  # - Use an SDK
  #
  # You should log the response code/output.

  log "Restarting instance now: ${instance_id}"

  # Placeholder command (will fail unless replaced):
  # hcloud ecs restart-server --instance-id "$instance_id" --region "$REGION"

  # Simulate success for template purposes:
  sleep 1

  log "Restart command issued for instance: ${instance_id}"
}

# ============================
# Main
# ============================

mkdir -p "$(dirname "$LOG_FILE")" 2>/dev/null || true

total_instances="${#INSTANCE_IDS[@]}"
log "----- Restart run started. DRY_RUN=${DRY_RUN}, total_instances=${total_instances} -----"

# Acquire lock
if ! acquire_lock; then
  log "Another restart run appears to be in progress. Exiting." 
  exit 0
fi

# Ensure lock is released on exit
trap 'release_lock' EXIT

# Time window check
if ! in_time_window; then
  log "Current time is outside allowed window (${ALLOW_START_HOUR}:00-${ALLOW_END_HOUR}:59). Exiting."
  exit 0
fi

# Restart with concurrency limit
restarted_count=0

for instance_id in "${INSTANCE_IDS[@]}"; do
  # If we've reached max concurrent restarts, wait.
  if [[ $restarted_count -ge $MAX_CONCURRENT_RESTARTS ]]; then
    # Wait for background jobs
    wait -n || true
    restarted_count=$((restarted_count - 1))
  fi

  # Start restart in background to allow concurrency.
  (
    restart_instance "$instance_id"
  ) &

  restarted_count=$((restarted_count + 1))
done

# Wait for remaining background jobs
wait || true

log "----- Restart run finished -----"

That script is a “template with a spine.” It demonstrates safe structure. Replace the restart placeholder with real Huawei Cloud calls.

How to Replace the Placeholder Restart Call

Here’s the philosophy: you want a reliable “restart instance” function that returns success/failure clearly.

If You Use a Huawei Cloud CLI

Some teams use a CLI tool configured with credentials. In that case, your restart function might call a CLI subcommand like “ecs server restart.” The exact command name may vary by tool/version.

Best practice:

  • Capture exit codes
  • Log stdout/stderr
  • Fail fast if an instance restart call fails (or record failures and continue—your call)

In your function, you might do something like:

# Example pseudo-command (replace with real CLI syntax)
# if hcloud ecs restart-server --instance-id "$instance_id"; then
#   log "Restart succeeded: $instance_id"
# else
#   log "Restart failed: $instance_id"
#   return 1
# fi

If You Call the ECS Restart API Directly

Direct API calls generally require:

  • An authentication mechanism (token retrieval, or signing requests)
  • The correct API endpoint for ECS restart action
  • Request headers and parameters

Because the signing method and endpoints are easy to get subtly wrong, you should test with a single instance first and verify the response.

In the script, you’d typically use curl and check the HTTP status code and response JSON (or at least the presence of expected fields).

Scheduling with Cron

Once your script is ready, schedule it. Assume the script is located at /opt/scripts/huawei_ecs_scheduled_restart.sh.

Make it executable:

chmod +x /opt/scripts/huawei_ecs_scheduled_restart.sh

Then edit your crontab:

crontab -e

Add a line like:

# Run daily at 03:00
0 3 * * * /opt/scripts/huawei_ecs_scheduled_restart.sh

Important: cron runs in the system time zone. If you’re running in UTC but your business clock lives in local time, you’ll schedule restarts at the wrong hour. The universe loves comedy; don’t feed it.

Testing Plan: The “Don’t Blow Up Production” Checklist

Here’s a sane testing approach:

  • Start with DRY_RUN=1 and run the script manually. Confirm it logs the intended instances.
  • Huawei Cloud USDT Top-up Reduce the allowlist to a single non-critical instance.
  • Turn DRY_RUN=0 and run during a maintenance window.
  • Verify instance state changes after the restart request. Confirm the instance returns to “running” (or equivalent) and services recover.
  • Review logs and confirm failure cases are visible.

Adding Health Checks (Because “Restarted” Isn’t the Same as “Okay”)

A restart script doesn’t automatically guarantee application health. After a reboot, instances may take time to come back up, or services might not recover properly due to configuration, missing mounts, or database connectivity delays.

So you can enhance the script by adding health checks. Common patterns:

  • Service-level checks: call a local health endpoint, check systemd status, verify ports.
  • Load balancer checks: confirm targets are healthy.
  • Huawei Cloud USDT Top-up Cluster checks: ensure quorum before restarting more nodes.

For example, before restarting an instance, you might verify its workload is not critical, or that enough redundancy exists. After restart, you might wait and re-check health before proceeding.

If you’re operating a highly available service, you can restart instances in batches (for example, max 1 or max 2 at a time). That keeps overall capacity stable while you refresh nodes.

Logging and Observability: Give Yourself Receipts

Logging is more than “print stuff and hope.” You want logs that help you answer:

  • Which instances did we attempt to restart?
  • Were the API calls successful?
  • How long did each step take?
  • Did the script exit early due to time window or lock?

In your script, you already log start/finish and each action. Consider adding:

  • Per-instance result codes (success/failure, API error messages)
  • Elapsed time per instance restart call
  • Counts of successes and failures

If you want extra professionalism, you can emit structured logs (JSON lines). But even plain text is fine if it’s consistent and timestamped.

Failure Handling: What If One Instance Restart Fails?

Failure is not a moral failing of your script. It’s just reality: API rate limits, permissions, temporary network issues, or instance state mismatches happen.

You need a policy:

  • Fail fast: stop the entire run on the first failure.
  • Continue and record: try all instances, collect failures, then exit with non-zero code.
  • Huawei Cloud USDT Top-up Retry: retry transient failures (with backoff) a limited number of times.

The template currently attempts restarts with concurrency and doesn’t stop everything on a single failure (because background jobs). If you want strict behavior, you can modify the script to run sequentially, or capture exit codes from background jobs (slightly more code, slightly more control).

Security Considerations (The “Don’t Leak Your Credentials” Section)

Credentials are valuable. Treat them like they’re made of gold and placed inside a suspicious-looking vending machine.

Huawei Cloud USDT Top-up Best practices:

  • Do not hardcode secrets in the script.
  • Use environment variables provided by your system/CI.
  • Restrict file permissions for config files that contain tokens.
  • Use least-privilege access: the account should only have permissions necessary for restart actions.

Going Beyond: Tag-Based Selection and “Restart Groups”

Once the basic script works, you can evolve it. Tag-based selection can reduce manual updates. For instance:

  • You define a tag: restart-group = nightly
  • Your script queries instances with that tag
  • Your script filters further by environment (prod/stage) or owner
  • Your script restarts only instances in the correct group

This requires implementing a “list instances” function and then restarting those returned IDs. The same safety rules apply: dry run first, logging always, and strict filters to avoid surprise inclusions.

Example Enhancements You Might Want

Depending on your environment, you can add these features:

  • Weekly/monthly cadence based on day-of-week and day-of-month.
  • “Restart only if uptime exceeds X days” to avoid unnecessary reboots.
  • Batch restarts: restart in waves, and wait between waves until instances are healthy.
  • Notifications: send a message to Slack/Teams/email when restarts start or finish.
  • Auditing: store a record in a database or file so you know last restart time.

A Note on “Restart” Semantics

In ECS environments, “restart” can mean different things: a soft reboot versus a hard reset, or simply triggering a stop/start sequence. Different options have different behavior around:

  • Instance downtime
  • Persistence of ephemeral state
  • Impact on attached volumes and network

When you implement the restart call, double-check what the Huawei Cloud ECS restart operation does in your case. Also confirm whether it affects the instance’s network and public IP behavior (usually it’s stable, but configurations vary).

Putting It All Together: A Reliable Workflow

Here’s a full “operationally sane” workflow:

  1. Start with DRY_RUN=1 and log-only behavior.
  2. Use an allowlist of one or two safe instances.
  3. Huawei Cloud USDT Top-up Run the script manually and verify logs.
  4. Integrate real Huawei Cloud restart command/API in the restart_instance function.
  5. Run during a maintenance window with DRY_RUN=0.
  6. Verify instance recovery and application health.
  7. Enable cron scheduling.
  8. After a successful period (say a couple of runs), expand the instance list gradually.

FAQ

Can I run the script from one of the ECS instances being restarted?

You can, but it’s risky if the script runs on the same instance that you’re restarting. Prefer running the scheduler from a separate management host or CI runner.

What if the script runs but the API returns “instance in invalid state”?

That usually means the instance is already restarting, stopped, or otherwise not eligible. Log the response, and consider adding a pre-check step to ensure the instance is in a restartrable state.

Is a lock enough to prevent duplicate runs?

For most cases, yes. The lock prevents overlapping runs on the same host. If you have multiple schedulers or hosts, use a centralized locking mechanism (or ensure only one scheduler exists).

Should I restart all instances every time?

Unless you’re confident you have redundancy and your services tolerate churn, no. Batch restarts with a max concurrency limit are usually safer.

Final Thoughts (With Minimal Chaos, Please)

A Huawei Cloud ECS scheduled restart script can be a powerful operational tool—or a source of spontaneous fireworks. The difference is not luck. It’s structure: time windows, allowlists, dry runs, logging, concurrency limits, and careful replacement of the restart placeholder with the correct Huawei Cloud restart mechanism.

If you follow the design in this article, you’ll end up with a script that behaves predictably. And predictability is the closest thing operations engineers get to peace.

TelegramContact Us
CS ID
@cloudcup
TelegramSupport
CS ID
@yanhuacloud