Tencent Cloud Credit Voucher Top-up Tencent Cloud Cloud Monitor Alert Setup Tutorial
Why Cloud Monitor Alerts Matter
When something goes wrong in production—CPU spikes, disk fills up, a service stops responding—waiting for manual checks is slow and risky. Tencent Cloud Cloud Monitor (often referred to as Cloud Monitor) helps you detect abnormal conditions in near real time and notify the right people immediately.
A good alert setup is not just about turning on notifications. It’s about creating alert rules that match real business impact, using notification channels people actually read, and tuning thresholds so you avoid alert fatigue. This tutorial walks you through a practical, end-to-end workflow for setting up alerts with Tencent Cloud Cloud Monitor.
Prerequisites Before You Start
1) Know what you want to protect
Before creating any rule, decide what “bad” means for your system. Examples: API latency exceeds a certain value, instance CPU stays above a threshold for several minutes, or a load balancer returns too many errors.
Good alerts are usually tied to user experience or service stability. If an alert doesn’t connect to an action you can take, it’s hard to justify.
2) Make sure the resources can be monitored
Cloud Monitor typically relies on metrics collected from your resources. Depending on your setup, you may need to enable monitoring for:
- Cloud servers (CVM)
- Managed components (for example, databases, load balancers, middleware)
- Application metrics (if you export custom metrics)
If metrics don’t appear, alert rules won’t work. Spend a little time verifying that you see the relevant metrics in Cloud Monitor first.
3) Have a notification plan
Think about how notifications should reach your team. Common options include message channels (like SMS, email, or chat-based channels), webhook integrations, or incident management tools. Prepare at least one “primary” channel and one “backup” channel.
Navigate to Cloud Monitor and Find Alert Settings
Open the Tencent Cloud console, then go to the Cloud Monitor section. The interface may vary slightly depending on your console version, but you’ll generally find an entry for alerts or alarm policies.
Your setup typically involves three components:
- Alarm policy (alert rule): defines the metric, condition, threshold, and evaluation period.
- Notification channel: where the alert message will be sent.
- Notification or receivers: who receives the message and how they are grouped.
Step 1: Confirm Metrics and Time Range
Start by locating the metric
Before writing an alert rule, find the metric you want to monitor. Use the metrics explorer or the monitoring dashboard to check:
- Is the metric present?
- Is it updated at the expected frequency?
- Does it have a stable baseline?
For example, if you want to alert on CPU usage, confirm you can see CPU utilization over the last several hours or days.
Choose realistic thresholds using historical data
Alerts work best when thresholds reflect real behavior. Look at typical ranges during normal operations, plus known periods of load (like peak traffic). If you set thresholds too close to normal, you’ll get frequent false alarms.
If you’re unsure, start with a conservative threshold and adjust after a few weeks of observation.
Step 2: Create an Alarm Policy (Alert Rule)
Tencent Cloud Credit Voucher Top-up Once you have the metric and threshold idea, create an alarm policy. The key parts usually include: metric selection, condition, trigger logic, and evaluation window.
1) Select the monitoring object
Choose the resource scope for the alarm rule. Depending on the product, this could be a single instance, a group, or a collection of resources. Pick the scope that matches your operational responsibility.
If your team owns a specific group of instances, use that group scope. If you use too broad a scope, it becomes harder to determine what specifically is failing.
2) Choose the metric and statistic
Tencent Cloud Credit Voucher Top-up Metrics often have multiple statistic types such as:
- Average
- Maximum
- Minimum
- Sum or rate (depending on the metric)
Choose the statistic that fits the scenario. For CPU overheating protection, maximum or average may both make sense, but max can trigger on brief spikes. For stability, average over several minutes can be safer.
Tencent Cloud Credit Voucher Top-up 3) Define the condition and threshold
Common conditions are “greater than,” “less than,” or “outside a range.” Examples:
- CPU > 80%
- Free disk space < 10 GB
- Error rate > 2%
Tencent Cloud Credit Voucher Top-up Consider whether the metric unit is percentage, bytes, seconds, or requests per minute. A mismatch here is one of the most frequent causes of confusing alerts.
4) Set the evaluation period (how long to trigger)
Cloud Monitor often evaluates conditions over an interval. Instead of triggering instantly, you can require that the condition holds for a continuous period.
For example:
- Require 5 minutes of sustained high CPU before triggering.
- Trigger immediately for critical metrics like service down, if supported.
This reduces noise from short spikes.
5) Decide the trigger severity or level
If the system supports severity levels (often called levels such as “Warning,” “Critical,” etc.), map them to real response expectations:
- Warning: investigate, but no immediate escalation.
- Critical: immediate attention, escalate quickly.
Even if you don’t have a full incident management system yet, using levels helps you route alerts later.
Step 3: Add and Test Notification Channels
Tencent Cloud Credit Voucher Top-up An alert without a working notification route is essentially a dashboard-only reminder. Ensure your notification channel is configured correctly.
1) Create or select notification destinations
In the alarm policy editor, you typically choose:
- Channel type (email, SMS, webhook, chat, etc.)
- Receiver list (users or groups)
- Optional template settings (message content)
2) Use different channels for different severities
A practical approach:
- Warning: send to a team channel or email.
- Critical: send to SMS or a high-priority channel, plus a chat mention.
This prevents your phone from ringing for every minor issue while still ensuring you’re reachable for major events.
3) Test before relying on it
Many alert systems allow a test notification. Use it. Confirm that:
- The message arrives quickly.
- The content includes enough context (metric name, threshold, current value, instance).
- The receiver understands what to do next.
If the message lacks context, you may need to adjust message templates or add labels to the alert rule.
Step 4: Best-Practice Alert Patterns (What to Create First)
Starting with a solid baseline set of alerts will save you time. Here are patterns that are widely useful, written in a way you can adapt to your environment.
Pattern A: Resource saturation
Use alerts for CPU, memory, and disk saturation. These often predict performance degradation before users complain.
- CPU utilization high for 5–10 minutes
- Tencent Cloud Credit Voucher Top-up Memory available low (or memory usage high)
- Disk space low for sustained period
Make thresholds match your capacity. If you run instances with autoscaling, you still want early warnings so scaling can react in time.
Pattern B: Service health and error rates
If your platform provides request metrics, track error rates and availability.
- Error rate > 1% for 3–5 minutes
- HTTP 5xx count or ratio above threshold
- Availability drops below expected level
Prefer alert conditions that reflect user impact rather than internal noise.
Pattern C: Latency and performance degradation
Latency alerts should focus on tail behavior (if you have it) or a meaningful summary like average/95th percentile.
- p95 response time above X ms for 5 minutes
- Queue length or request processing time above threshold
Latency can spike briefly without major impact. That’s why an evaluation window matters.
Pattern D: Data freshness (for critical pipelines)
Tencent Cloud Credit Voucher Top-up If you run scheduled jobs or data pipelines, alert when outputs are delayed.
- Tencent Cloud Credit Voucher Top-up Job success rate drops
- Lag between event time and processed time increases
- Pipeline stops producing data
These alerts are often essential for analytics or downstream systems.
Step 5: Tune to Reduce False Alarms
The first version of alert rules is rarely perfect. Most teams need at least one tuning cycle.
1) Adjust thresholds based on real history
If you receive alerts during normal usage, raise the threshold or extend the evaluation period.
If you miss incidents that you care about, lower the threshold or shorten the evaluation period (but be careful not to create noise).
2) Add cool-down or repeat suppression (if available)
Many systems allow you to avoid repeated notifications while the condition keeps failing. Use a cool-down interval so that a single incident doesn’t spam your channels.
For example, notify once every 10–30 minutes while the issue is unresolved.
3) Avoid alerting on metrics with unstable collection
If a metric sometimes disappears or updates irregularly, your alert may trigger unexpectedly. Validate that metrics are reliable before building critical rules on them.
4) Consider maintenance windows
During deployments, scaling, or planned maintenance, alerts may not be meaningful. If your console supports it, configure a quiet period or maintenance window for those times.
Step 6: Operational Workflow After an Alert Triggers
A good alert is part of an operational loop. When a notification arrives, you should have a repeatable workflow.
1) Read the message fully
Before opening dashboards, check:
- Which resource is affected
- What metric triggered
- Current value, threshold, and evaluation period
- Alert time and severity
2) Confirm it’s real in the dashboard
Open the metric chart around the alert time. Look for:
- Whether the condition sustained long enough to trigger
- Whether the metric is behaving normally after the alert
Sometimes you’ll see a brief spike that matches the evaluation window poorly. That’s feedback for tuning.
3) Link symptoms to likely causes
Resource saturation alerts can help you narrow down. For example:
- High CPU with increased latency suggests compute bottleneck.
- Low disk with errors can indicate storage exhaustion.
- High error rate without CPU increase may point to dependency failures.
Even if you don’t have a full runbook, your team should document the most common correlations.
4) Decide whether to escalate
Use severity levels and your team’s response policy. Warning alerts might be investigated within business hours; critical alerts might require immediate action and incident escalation.
Example: A Complete Setup for a Typical Web Service
Let’s put it together with a simple scenario: you run a web service on multiple instances behind a load balancer.
1) Metrics to monitor
- CPU utilization on instances
- Memory utilization or available memory
- Disk free space
- HTTP 5xx error rate or error count
- Request latency (average or p95)
2) Suggested alert rules to start
- Warning CPU: CPU > 70% for 5 minutes
- Tencent Cloud Credit Voucher Top-up Critical CPU: CPU > 85% for 5 minutes
- Critical Disk: free disk space < 10 GB for 10 minutes
- Critical Error rate: HTTP 5xx ratio > 1% for 3 minutes
- Warning Latency: p95 latency > 300 ms for 5 minutes
- Critical Latency: p95 latency > 600 ms for 5 minutes
3) Notification routing
- Tencent Cloud Credit Voucher Top-up Warning alerts go to an operations chat channel and email list.
- Critical alerts go to SMS or high-priority chat mentions, plus email.
Then test each alert once to confirm the message includes the right context.
Common Mistakes to Avoid
- Using one-size-fits-all thresholds: different services have different baselines.
- Forgetting evaluation windows: instant triggers cause noisy alerts.
- Not testing notification delivery: alerts might fail silently.
- Too many alerts at the start: focus on the top risks first.
- Ignoring alert message quality: if recipients don’t get context, they waste time.
Checklist: Your Alert Setup Is Done When...
- You can see the metrics you plan to alert on.
- Each alarm rule has a clear threshold, statistic type, and evaluation period.
- Notification channels are configured and tested.
- Warning and critical levels are routed differently.
- You have a simple response workflow for common alert types.
- You reviewed alert frequency and tuned it after initial runs.
Next Steps
After you finish the first alert set, the improvement work is ongoing. As your system changes, update thresholds, add new metrics tied to business impact, and refine your notification routing. Over time, your alerts will become a dependable early warning system rather than a source of noise.
If you tell me what resources you’re monitoring (CVM, load balancer, database, Kubernetes, or custom apps) and what you want to catch first (CPU, latency, errors, or data pipeline delays), I can suggest specific alert rules and a clean routing plan for your team.

