{"id":1182,"date":"2026-06-18T23:31:00","date_gmt":"2026-06-18T23:31:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/"},"modified":"2026-06-19T00:59:24","modified_gmt":"2026-06-19T00:59:24","slug":"monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/","title":{"rendered":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div id=\"\">\n<p>Monitoring and troubleshooting generative AI inference endpoints working at scale is difficult. When your massive language mannequin (LLM) endpoint\u2019s P99 latency spikes, you will need to decide in minutes whether or not the basis trigger is GPU reminiscence stress, a saturated KV cache, unbalanced visitors throughout Availability Zones, or an auto scaling coverage that hasn\u2019t triggered. The shift from coaching to serving is reshaping how groups deploy LLMs and different generative AI fashions in manufacturing. Machine studying (ML) platform engineers, MLOps groups, and website reliability engineers (SREs) should maintain inference endpoints wholesome, responsive, and cost-efficient, usually throughout dozens of fashions and tons of of GPU cases.<\/p>\n<p>Amazon SageMaker AI gives totally managed real-time inference internet hosting for machine studying fashions. You deploy a mannequin to a SageMaker endpoint backed by a number of compute cases, and SageMaker handles provisioning and scaling. SageMaker helps a number of endpoint architectures. This put up focuses on the 2 most related to generative AI workloads with detailed observability:<\/p>\n<p>        Single-model endpoints (SME) \u2013 Every endpoint hosts one mannequin on devoted cases. SMEs are easy to arrange and purpose about, however every mannequin requires its personal fleet of GPU cases.<br \/>\n        Inference part (IC) endpoints \u2013 A number of fashions share the identical set of cases by way of inference elements. Every inference part defines a mannequin, its useful resource necessities (CPU, GPU, reminiscence), and its scaling coverage. IC endpoints are the beneficial structure for manufacturing generative AI workloads as a result of they assist multi-model internet hosting on shared GPU infrastructure, impartial scaling per mannequin, and excessive availability (HA) by way of copy distribution throughout AZs.<\/p>\n<p>SageMaker endpoints emit metrics like invocation counts, mannequin latency, and overhead latency to Amazon CloudWatch. These combination metrics are helpful for understanding total endpoint well being. As a result of groups scale to multi-model deployments on GPU fleets, they want deeper alerts. Amazon SageMaker AI now emits over 100 detailed inference metrics. These cowl GPU well being, token-level latency, KV cache stress, visitors distribution throughout AZs, inference part placement, and chilly begin diagnostics. These metrics circulate to a built-in SageMaker Insights dashboard in Amazon CloudWatch, a completely managed observability resolution that removes the necessity for customized Grafana dashboards and Prometheus configuration. The SageMaker Insights dashboard helps each endpoint sorts and routinely reveals IC-specific panels when inference elements are detected.<\/p>\n<p>For extra particulars on SageMaker inference, see Deploy fashions for real-time inference.<\/p>\n<p>On this put up, you&#8217;ll discover ways to:<\/p>\n<p>        Activate detailed observability metrics on new and present SageMaker inference endpoints.<br \/>\n        Navigate the SageMaker Insights dashboard to observe fleet well being throughout Efficiency, Capability, and Reliability views.<br \/>\n        Join the metrics to your personal observability device (Grafana, Datadog) by way of the PromQL-compatible endpoint.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-1.png\" alt=\"Architecture diagram of SageMaker inference endpoints emitting OpenTelemetry metrics to Amazon CloudWatch and the SageMaker Insights dashboard\" width=\"600\"\/><\/p>\n<h2 id=\"sagemaker-inference-observability-overview\">SageMaker inference observability overview<\/h2>\n<p>SageMaker inference endpoints emit native OpenTelemetry metrics to CloudWatch. The SageMaker Insights dashboard is situated within the CloudWatch console beneath Infrastructure Monitoring \u2192 SageMaker Insights. It queries these metrics utilizing PromQL and renders visualizations on the fleet, endpoint, and inference-component stage throughout three tabs: Efficiency, Capability, and Reliability.<\/p>\n<p>        Efficiency \u2013 Fleet well being, token latency, throughput, errors, engine stress.<br \/>\n        Capability \u2013 GPU, CPU, and reminiscence utilization of the fleet.<br \/>\n        Reliability \u2013 Availability Zone distribution, scaling occasions, chilly begin anatomy, and inadequate capability errors.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-2.png\" alt=\"SageMaker Insights dashboard in CloudWatch showing the Performance, Capacity, and Reliability tabs\" width=\"600\"\/><\/p>\n<h3 id=\"key-services\">Key companies<\/h3>\n<p>        Amazon SageMaker AI \u2013 Managed inference with endpoints and inference elements.<br \/>\n        Amazon CloudWatch \u2013 Native assist for OpenTelemetry metrics and PromQL queries by way of SageMaker Insights.<\/p>\n<p>For background on the OpenTelemetry and PromQL assist in CloudWatch, see Introducing OpenTelemetry PromQL assist in Amazon CloudWatch.<\/p>\n<h2 id=\"prerequisites\">Stipulations<\/h2>\n<p>It&#8217;s essential to have the next to observe together with this put up.<\/p>\n<p>        An AWS account with a minimum of one SageMaker real-time inference endpoint.<br \/>\n        AWS Id and Entry Administration (IAM) permissions: sagemaker:CreateEndpointConfig, sagemaker:UpdateEndpoint, and cloudwatch:GetMetricData.<br \/>\n        vLLM or SGLang container framework (required for token-level metrics like TTFT and ITL).<\/p>\n<p>GPU cases obtain per-accelerator utilization metrics along with the CPU and reminiscence metrics accessible on all occasion sorts. For the complete setup information, see Getting began with detailed observability.<\/p>\n<h2 id=\"activate-detailed-metrics-on-your-endpoints\">Activate detailed metrics in your endpoints<\/h2>\n<h3 id=\"new-endpoints-automatic-default-on\">New endpoints: Automated (default-on)<\/h3>\n<p>For any new endpoint configurations you create, detailed metrics are turned on by default. The EnableDetailedObservability parameter in your endpoint configuration defaults to true. No further code is required.<\/p>\n<div class=\"hide-language\">\n        import boto3<\/p>\n<p>sm = boto3.shopper(&#8220;sagemaker&#8221;)<\/p>\n<p># Create endpoint config \u2014 observability turned on by default<br \/>\nresponse = sm.create_endpoint_config(<br \/>\n    EndpointConfigName=&#8221;my-llm-config&#8221;,<br \/>\n    ProductionVariants=[{<br \/>\n        &#8220;VariantName&#8221;: &#8220;primary&#8221;,<br \/>\n        &#8220;InstanceType&#8221;: &#8220;ml.g6.4xlarge&#8221;,<br \/>\n        &#8220;InitialInstanceCount&#8221;: 2,<br \/>\n        &#8220;ManagedInstanceScaling&#8221;: {<br \/>\n            &#8220;Status&#8221;: &#8220;ENABLED&#8221;,<br \/>\n            &#8220;MinInstanceCount&#8221;: 2,<br \/>\n            &#8220;MaxInstanceCount&#8221;: 8<br \/>\n        }<br \/>\n    }],<br \/>\n    ExecutionRoleArn=&#8221;arn:aws:iam::123456789012:position\/SageMakerExecutionRole&#8221;\n       <\/p><\/div>\n<p>The EnableDetailedObservability flag in your endpoint configuration defaults to true, so no further configuration is required. It&#8217;s also possible to explicitly set the publishing frequency utilizing MetricsPublishFrequencyInSeconds in MetricsConfig. The default is 60 seconds. For workloads that want close to real-time monitoring, you possibly can set it to lower than a minute.<\/p>\n<div class=\"hide-language\">\n        # Create endpoint<br \/>\nsm.create_endpoint(<br \/>\n    EndpointName=&#8221;my-llm-endpoint&#8221;,<br \/>\n    EndpointConfigName=&#8221;my-llm-config&#8221;<br \/>\n)\n       <\/div>\n<p>Inside 2 minutes of the endpoint reaching InService, the OpenTelemetry format metrics start flowing to CloudWatch.<\/p>\n<h3 id=\"existing-endpoints-opt-in\">Present endpoints: Decide-in<\/h3>\n<p>Present endpoints require an specific opt-in. Create a brand new endpoint configuration with the MetricsConfig flag, then replace your endpoint. This follows the identical sample as any endpoint configuration change.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-3.png\" alt=\"SageMaker console showing the Enable detailed observability option in endpoint configuration\" width=\"600\"\/><\/p>\n<div class=\"hide-language\">\n        # Step 1: Create new config with detailed observability turned on<br \/>\nsm.create_endpoint_config(<br \/>\n    EndpointConfigName=&#8221;my-existing-config-v2&#8243;,<br \/>\n    ProductionVariants=[{<br \/>\n        &#8220;VariantName&#8221;: &#8220;primary&#8221;,<br \/>\n        &#8220;ModelName&#8221;: &#8220;my-existing-model&#8221;,<br \/>\n        &#8220;InstanceType&#8221;: &#8220;ml.g6.4xlarge&#8221;,<br \/>\n        &#8220;InitialInstanceCount&#8221;: 2<br \/>\n    }],<br \/>\n    MetricsConfig={&#8220;EnableDetailedObservability&#8221;: True},<br \/>\n    ExecutionRoleArn=&#8221;arn:aws:iam::123456789012:position\/SageMakerExecutionRole&#8221;<br \/>\n)<\/p>\n<p># Step 2: Replace endpoint<br \/>\nsm.update_endpoint(<br \/>\n    EndpointName=&#8221;my-existing-endpoint&#8221;,<br \/>\n    EndpointConfigName=&#8221;my-existing-config-v2&#8243;<br \/>\n)\n       <\/p><\/div>\n<p>The SageMaker console additionally gives a guided three-step wizard after you select Allow detailed observability: study concerning the metrics, activate OTel enrichment, and choose which endpoints to choose in.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-4.png\" alt=\"Three-step wizard in the SageMaker console for enabling detailed observability on existing endpoints\" width=\"600\"\/><\/p>\n<h3 id=\"enable-otel-enrichment-for-classic-cloudwatch-metrics\">Allow OTel enrichment for traditional CloudWatch metrics<\/h3>\n<p>Native OpenTelemetry metrics circulate routinely to CloudWatch after enablement. Nonetheless, present basic metrics (Invocations, ModelLatency, OverheadLatency) require OTel enrichment to be seen within the SageMaker Insights dashboard and queryable with PromQL.<\/p>\n<p>Navigate to CloudWatch Console then Settings and activate OTel metric enrichment and Useful resource tags for telemetry. This can be a one-time, account-level and AWS Area-level setting.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-5.png\" alt=\"CloudWatch Settings page with OTel metric enrichment and Resource tags for telemetry options selected\" width=\"600\"\/><\/p>\n<h2 id=\"navigate-to-the-sagemaker-insights-dashboard-from-the-sagemaker-console\">Navigate to the SageMaker Insights dashboard from the SageMaker console<\/h2>\n<p>You may entry the SageMaker Insights dashboard by way of both the SageMaker console or the CloudWatch console. Inside SageMaker, there are three entry factors, every pre-filtered to their context:<\/p>\n<p>          #<br \/>\n          Entry Level<br \/>\n          Filter Utilized<br \/>\n          Use Case<\/p>\n<p>          1<br \/>\n          Endpoints listing web page \u2192 \u201cOpen SageMaker Insights\u201d<br \/>\n          Fleet-level (all endpoints)<br \/>\n          \u201cGive me the large image\u201d<\/p>\n<p>          2<br \/>\n          Endpoint element web page \u2192 \u201cView in SageMaker Insights\u201d<br \/>\n          Filtered to that endpoint<br \/>\n          \u201cDrill into this particular endpoint\u201d<\/p>\n<p>          3<br \/>\n          IC tab \u2192 per-IC \u201cMetrics\u201d hyperlink<br \/>\n          Filtered to endpoint + IC<br \/>\n          \u201cDebug this inference part\u201d<\/p>\n<p>Each path deep-links with pre-applied filters, so that you received\u2019t land on a clean dashboard looking for your sources.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-6.png\" alt=\"SageMaker console with three deep-link entry points to the SageMaker Insights dashboard\" width=\"600\"\/><\/p>\n<h2 id=\"performance-tab-monitoring-fleet-health-and-debugging-latency\">Efficiency tab: Monitoring fleet well being and debugging latency<\/h2>\n<p>The Efficiency tab is the place most prospects spend their time. It solutions questions like \u201cIs the whole lot operating nicely?\u201d and \u201cIf not, which part is the issue?\u201d The Efficiency tab contains a number of time-series panels that work collectively to pinpoint latency points.<\/p>\n<h3 id=\"performance-health-and-instance-performance-table\">Efficiency well being and occasion efficiency desk<\/h3>\n<p>Coloration-coded hexagons visualize each useful resource in your fleet. Toggle between Situations, IC Copies, and Endpoints views. The hexagon shade signifies state:<\/p>\n<p>        Inexperienced for OK.<br \/>\n        White for no alarms detected.<br \/>\n        Pink for in alarm.<\/p>\n<p>Hover over any hexagon to see occasion sort, TTFT, output TPS, concurrent requests, KV cache utilization, and CloudWatch alarm standing. Select Filter by this occasion to drill down. Each panel on the web page updates to indicate solely that occasion\u2019s knowledge.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-8.png\" alt=\"Honeycomb hexagon visualization with a hover card showing per-instance performance metrics\" width=\"600\"\/><\/p>\n<p>The desk reveals each occasion with efficiency metrics side-by-side. Use this desk to identify outliers in TTFT, output TPS, and concurrent requests. The TTFT, Output TPS, Concurrent Requests, and KV Cache columns present knowledge emitted by the vLLM and SGLang frameworks solely.<\/p>\n<p>The Token streaming panel plots Time to First Token (TTFT) and Inter-Token Latency (ITL) over time with a P50\/P99 toggle. TTFT measures how lengthy customers wait earlier than seeing the primary response character. ITL measures time between consecutive tokens, which instantly impacts streaming smoothness. You may filter by endpoint, inference part title, or mannequin to isolate which part contributes to latency.<\/p>\n<p>Once you establish a TTFT spike, the Latency breakdown panel helps you attribute it. This panel separates complete latency into Mannequin Latency (time the mannequin spends processing) and Overhead Latency (time the platform spends routing and scheduling). An Invoke tab reveals the complete request path, and a Streaming tab reveals time-to-first-chunk particularly. If each Mannequin Latency and Overhead Latency are regular however TTFT remains to be elevated, the mannequin\u2019s inference engine could be holding requests in its inner queue, for instance, ready for KV cache slots. Examine the Engine and request stress panel to verify.<\/p>\n<p>The Site visitors distribution panel reveals per-instance or per-inference-component request circulate with Availability Zone filtering. Toggle the AZ dropdown to isolate visitors by zone. If one AZ reveals zero visitors whereas others are loaded, that signifies a routing or placement concern. You should use the occasion\/IC toggle to change between \u201cWhich machines deal with visitors?\u201d and \u201cWhich fashions deal with visitors?\u201d views.<\/p>\n<p>Lastly, the Token throughput panel measures precise tokens processed per second, damaged down by enter\/output, percentiles, or by occasion. This instantly measures inference effectivity. For instance, in case your ml.g6.4xlarge delivers 150 tokens per second output when the mannequin benchmark reveals 500, that signifies a useful resource constraint, configuration concern, or KV cache stress. The multi-framework legend (SGLang, vLLM, DJL) lets multi-model endpoints evaluate throughput throughout inference engines.<\/p>\n<h3 id=\"engine-and-request-pressure\">Engine and request stress<\/h3>\n<p>The Engine and request stress panel is your early warning system for stopping outages.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-9.png\" alt=\"Engine and request pressure panel showing KV cache utilization, running requests, and waiting requests over time\" width=\"600\"\/><\/p>\n<p>The time-series view reveals the per-framework breakdown, with tooltips that present actual values at any timestamp. In the event you see KV cache repeatedly climbing to 40\u201350 % throughout enterprise hours, configure autoscaling to set off at a threshold worth earlier than prospects really feel the impression.<\/p>\n<h2 id=\"capacity-tab-planning-deployments-and-resource-management\">Capability tab: Planning deployments and useful resource administration<\/h2>\n<p>The Capability tab solutions questions like \u201cDo I&#8217;ve sufficient sources?\u201d, \u201cThe place is there headroom?\u201d, and \u201cCan I match one other mannequin?\u201d<\/p>\n<h3 id=\"capacity-health\">Capability well being<\/h3>\n<p>The identical honeycomb visualization from Efficiency reappears right here, with useful resource utilization percentages within the hover card: GPU, GPU reminiscence, CPU, CPU reminiscence, and Disk.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-10.png\" alt=\"Capacity health honeycomb view with a hover card showing GPU, GPU memory, CPU, CPU memory, and disk utilization\" width=\"600\"\/><\/p>\n<p>Earlier than you deploy a brand new mannequin or scale copies, hover over cases in your goal endpoint. If GPU reminiscence is at 89 %, there\u2019s restricted VRAM headroom for added mannequin weights.<\/p>\n<h3 id=\"fleet-utilization-over-time\">Fleet utilization over time<\/h3>\n<p>This panel reveals useful resource consumption traits with toggles for Occasion, IC copies, and Endpoint aggregation. Key alerts embody the next:<\/p>\n<p>        GPU Reminiscence trending upward over days signifies that you just\u2019re approaching capability limits. Add cases earlier than utilization reaches the restrict.<br \/>\n        GPU Reminiscence dropping abruptly signifies {that a} mannequin crashed or was unloaded. Examine.<br \/>\n        Disk spikes that recur periodically correlate with mannequin downloads throughout chilly begins.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-11.png\" alt=\"Fleet utilization time series showing GPU, GPU memory, CPU, memory, and disk consumption with Instance, IC copies, and Endpoint toggles\" width=\"600\"\/><\/p>\n<h2 id=\"reliability-tab-supporting-high-availability-and-resilience-view\">Reliability tab: Supporting excessive availability and resilience view<\/h2>\n<p>The Reliability tab solutions questions like \u201cIf an AZ goes down, will my inference fleet survive?\u201d, \u201cAre scaling occasions working?\u201d, and \u201cWhy are chilly begins sluggish?\u201d<\/p>\n<h3 id=\"availability-zone-distribution\">Availability Zone distribution<\/h3>\n<p>A bar chart reveals occasion and IC copy counts per AZ. This view reveals your excessive availability posture.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-12.png\" alt=\"Bar chart of instance and IC copy counts per Availability Zone with Instances and IC Copies toggle\" width=\"600\"\/><\/p>\n<p>          Distribution<br \/>\n          Threat<br \/>\n          Motion<\/p>\n<p>          Even throughout over 3 AZs<br \/>\n          Low<br \/>\n          No motion<\/p>\n<p>          Concentrated in 1-2 AZs<br \/>\n          Medium<br \/>\n          Rebalance<\/p>\n<p>          0 cases in any AZ<br \/>\n          Excessive<br \/>\n          Single AZ failure takes you offline<\/p>\n<p>Toggle between Situations and IC Copies. Situations could be balanced, however IC copies could possibly be targeting just a few machines.<\/p>\n<h3 id=\"cold-start-anatomy\">Chilly begin anatomy<\/h3>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-13.png\" alt=\"Stacked bar chart breaking down each IC provisioning event into model download, GPU load, container start, and health check phases\" width=\"600\"\/><\/p>\n<p>Each IC provisioning occasion displayed as a horizontal stacked bar with 4 phases:<\/p>\n<p>          Part<br \/>\n          Coloration<br \/>\n          What it measures<br \/>\n          Optimization<\/p>\n<p>          Mannequin obtain<br \/>\n          Blue<br \/>\n          Pull mannequin weights from Amazon Easy Storage Service (Amazon S3)<br \/>\n          Compress artifacts, use Amazon Elastic File System (Amazon EFS) caching<\/p>\n<p>          GPU load<br \/>\n          Purple<br \/>\n          Load weights onto GPU<br \/>\n          Smaller quantization, pre-warming<\/p>\n<p>          Container begin<br \/>\n          Orange<br \/>\n          Container initialization<br \/>\n          Cut back dependencies<\/p>\n<p>Within the screenshot, gma-ic-vllm took 237.6 seconds, with mannequin obtain dominating, whereas gma-rblk-ic-tiny was solely 41.4 seconds as a result of it\u2019s a smaller mannequin. This view tells you which ones section to optimize for quicker scaling response occasions.<\/p>\n<h3 id=\"ice-diagnostics\">ICE diagnostics<\/h3>\n<p>The ICE diagnostics view tracks inadequate capability errors (ICE), which happen when SageMaker can\u2019t provision requested cases. The desk reveals:<\/p>\n<p>        When the failure occurred.<br \/>\n        Which endpoint was affected (deep-links to the console).<br \/>\n        Which occasion sort was unavailable.<br \/>\n        Which AZ had no capability.<\/p>\n<p>Within the previous screenshot, all 12 ICE occasions are for p5.48xlarge throughout all 4 AZs, indicating full regional exhaustion for this occasion sort. You now know to change to different occasion sorts as a fallback.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-14.png\" alt=\"ICE diagnostics table listing the time, affected endpoint, instance type, and Availability Zone for each insufficient capacity event\" width=\"600\"\/><\/p>\n<p>For groups with present Grafana or different PromQL-compatible instruments, you possibly can question SageMaker Insights metrics instantly out of your platform with out switching to the CloudWatch console. The next walkthrough demonstrates the setup utilizing Grafana. The identical steps apply to self-hosted Grafana or different suitable instruments, with minor configuration variations.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-15.png\" alt=\"SageMaker Insights metrics flowing through the PromQL endpoint to a Grafana dashboard\" width=\"600\"\/><\/p>\n<h3 id=\"step-1-get-the-promql-endpoint-url\">Step 1: Get the PromQL endpoint URL<\/h3>\n<p>Navigate to SageMaker Console, then choose Endpoints. From there, choose your endpoint after which select Connect with your observability device. Copy the displayed endpoint URL. It follows the format proven within the SageMaker console.<\/p>\n<h3 id=\"step-2-configure-your-grafana-data-source\">Step 2: Configure your Grafana knowledge supply<\/h3>\n<p>In Amazon Managed Grafana (Traditional CloudWatch 2.4+) or self-hosted Grafana with the Amazon Managed Service for Prometheus plugin (v3.0.0+):<\/p>\n<p>        Navigate to Configuration, Knowledge Sources, then Add knowledge supply. Choose Amazon Managed Service for Prometheus and set the URL to the PromQL endpoint URL from Step 1.<br \/>\n        Below Service Supplier, enter monitoring.<br \/>\n        Configure SigV4 authentication with an IAM position that has the cloudwatch:GetMetricData and cloudwatch:ListMetrics permissions.<br \/>\n        Select Save &amp; Check. It is best to see Knowledge supply is working.<\/p>\n<h3 id=\"step-3-import-the-pre-built-dashboard-template\">Step 3: Import the pre-built dashboard template<\/h3>\n<p>Obtain the dashboard template JSON from the identical Connect with your observability device web page within the SageMaker console. Import the downloaded JSON template into Grafana (Dashboards \u2192 Import), choose the Prometheus knowledge supply you configured in Step 2, and also you get pre-configured Efficiency, Capability, and Reliability panels matching the SageMaker Insights format.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-16.png\" alt=\"Imported Grafana dashboard with pre-configured Performance, Capacity, and Reliability panels matching SageMaker Insights\" width=\"600\"\/><\/p>\n<h3 id=\"step-4-query-metrics-with-promql\">Step 4: Question metrics with PromQL<\/h3>\n<p>With the info supply related, you possibly can write customized PromQL queries. For instance:<\/p>\n<p>       KV cache<br \/>\nvllm:kv_cache_usage_perc{&#8220;aws.sagemaker.endpoint.title&#8221;=&#8221;ep-prsn-ic&#8221;,&#8221;aws.sagemaker.inference_component.title&#8221;=&#8221;ic-qwen3-4b&#8221;}<\/p>\n<p># Energetic requests<br \/>\nvllm:num_requests_running{&#8220;aws.sagemaker.endpoint.title&#8221;=&#8221;ep-prsn-ic&#8221;,&#8221;aws.sagemaker.inference_component.title&#8221;=&#8221;ic-qwen3-4b&#8221;}<\/p>\n<p># TTFT P99<br \/>\nhistogram_quantile(0.99, fee(vllm:time_to_first_token_seconds{&#8220;aws.sagemaker.endpoint.title&#8221;=&#8221;ep-prsn-ic&#8221;,&#8221;aws.sagemaker.inference_component.title&#8221;=&#8221;ic-qwen3-4b&#8221;}[5m]))<\/p>\n<h2 id=\"pricing\">Pricing<\/h2>\n<p>SageMaker doesn\u2019t cost individually for emitting detailed observability metrics. The metrics are printed to Amazon CloudWatch in OpenTelemetry knowledge format, and customary CloudWatch OpenTelemetry ingestion pricing applies. OpenTelemetry metrics ingested into CloudWatch are charged at $0.50 per GB ingested. In the event you activate OTel vended metric enrichment (required to view basic CloudWatch metrics like Invocations and ModelLatency within the Insights dashboard), enriched metrics are additionally charged at $0.50 per GB. For detailed pricing examples and a price calculator, see the OpenTelemetry Metrics part on the Amazon CloudWatch pricing web page.<\/p>\n<h2 id=\"clean-up\">Clear up<\/h2>\n<p>To keep away from ongoing expenses, delete take a look at sources on this order:<\/p>\n<div class=\"hide-language\">\n        # Delete inference elements first (if IC endpoint)<br \/>\naws sagemaker delete-inference-component &#8211;inference-component-name my-ic<\/p>\n<p># Delete endpoints<br \/>\naws sagemaker delete-endpoint &#8211;endpoint-name my-endpoint<\/p>\n<p># Look forward to deletion, then delete configs<br \/>\naws sagemaker delete-endpoint-config &#8211;endpoint-config-name my-config\n       <\/p><\/div>\n<p>GPU cases are billed per second whereas endpoints are InService. Delete promptly after testing.<\/p>\n<h2 id=\"conclusion\">Conclusion<\/h2>\n<p>On this put up, you enabled SageMaker detailed metrics on inference endpoints and used the built-in SageMaker Insights dashboard to observe fleet well being, debug latency utilizing token-level metrics, validate excessive availability, and plan capability for brand new deployments.<\/p>\n<p>To get began, see the next sources:<\/p>\n<h2 id=\"acknowledgments\">Acknowledgments<\/h2>\n<p>The SageMaker Insights dashboard and detailed observability metrics are the results of shut collaboration between the Amazon SageMaker AI and Amazon CloudWatch groups. We thank the engineering, product, and options structure groups whose work made this launch potential.<\/p>\n<p>We additionally thank the next contributors for his or her overview and inputs on this weblog put up:<\/p>\n<p>        Felipe Lopez \u2013 Principal GenAI\/ML Architect, AWS<br \/>\n        Sandeep Raveesh-Babu \u2013 Sr.\u00a0Worldwide Specialist SA, GenAI, AWS<br \/>\n        Johna Liu \u2013 Sr.\u00a0Software program Growth Engineer, Amazon SageMaker<br \/>\n        Raviprakash Darbha \u2013 Sr.\u00a0Software program Growth Engineer, Amazon SageMaker<br \/>\n        Prajwal Kammardi \u2013 Software program Growth Engineer, Amazon SageMaker<br \/>\n        Jiaxi Xu \u2013 Software program Growth Engineer, Amazon SageMaker<br \/>\n        Orcun Berkem \u2013 Principal Engineer, Observability, Amazon CloudWatch<br \/>\n        Steve McCurry \u2013 Principal Product Supervisor, Amazon CloudWatch<\/p>\n<h2>Concerning the creator<\/h2>\n<div class=\"blog-author-box\">\n<div class=\"blog-author-image\">\n<p><img decoding=\"async\" loading=\"lazy\" class=\"alignleft size-full\" src=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ML-21272-17.png\" alt=\"Apoorva Chandra\" width=\"100\" height=\"100\"\/><\/p>\n<\/p><\/div>\n<h3 class=\"lb-h4\">Apoorva Chandra<\/h3>\n<p>Apoorva is a Senior Product Supervisor on the Amazon SageMaker AI Inference staff at AWS. She leads the inference observability initiative, targeted on serving to ML platform groups monitor, debug, and optimize GenAI workloads in manufacturing. Previous to SageMaker, she labored with the industrial utility staff on AWS for modernizing enterprise SAP prospects\u2019 journey on AWS. Exterior of labor, Apoorva enjoys mountaineering, exploring espresso retailers, and spending time with associates.<\/p>\n<\/p><\/div><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/aws.amazon.com\/blogs\/machine-learning\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Monitoring and troubleshooting generative AI inference endpoints working at scale is difficult. When your massive language mannequin (LLM) endpoint\u2019s P99 latency spikes, you will need to decide in minutes whether or not the basis trigger is GPU reminiscence stress, a saturated KV cache, unbalanced visitors throughout Availability Zones, or an auto scaling coverage that hasn\u2019t [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1184,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[7],"tags":[1591,1590,1585,1587,1586,1068,1589,1588,1584,84],"class_list":["post-1182","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-science-mlops","tag-cloudwatch","tag-dashboard","tag-debug","tag-detailed","tag-generative","tag-inference","tag-insights","tag-metrics","tag-monitor","tag-sagemaker"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Monitoring and troubleshooting generative AI inference endpoints working at scale is difficult. When your massive language mannequin (LLM) endpoint\u2019s P99 latency spikes, you will need to decide in minutes whether or not the basis trigger is GPU reminiscence stress, a saturated KV cache, unbalanced visitors throughout Availability Zones, or an auto scaling coverage that hasn\u2019t [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-18T23:31:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-19T00:59:24+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png\" \/><meta property=\"og:image\" content=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch\",\"datePublished\":\"2026-06-18T23:31:00+00:00\",\"dateModified\":\"2026-06-19T00:59:24+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/\"},\"wordCount\":2807,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/d2908q01vomqb2.cloudfront.net\\\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\\\/2026\\\/06\\\/18\\\/ml-21272.png\",\"keywords\":[\"CloudWatch\",\"dashboard\",\"debug\",\"detailed\",\"generative\",\"inference\",\"Insights\",\"metrics\",\"Monitor\",\"SageMaker\"],\"articleSection\":[\"Data Science &amp; MLOps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/\",\"name\":\"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/d2908q01vomqb2.cloudfront.net\\\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\\\/2026\\\/06\\\/18\\\/ml-21272.png\",\"datePublished\":\"2026-06-18T23:31:00+00:00\",\"dateModified\":\"2026-06-19T00:59:24+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#primaryimage\",\"url\":\"https:\\\/\\\/d2908q01vomqb2.cloudfront.net\\\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\\\/2026\\\/06\\\/18\\\/ml-21272.png\",\"contentUrl\":\"https:\\\/\\\/d2908q01vomqb2.cloudfront.net\\\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\\\/2026\\\/06\\\/18\\\/ml-21272.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/18\\\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/","og_locale":"en_US","og_type":"article","og_title":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24","og_description":"Monitoring and troubleshooting generative AI inference endpoints working at scale is difficult. When your massive language mannequin (LLM) endpoint\u2019s P99 latency spikes, you will need to decide in minutes whether or not the basis trigger is GPU reminiscence stress, a saturated KV cache, unbalanced visitors throughout Availability Zones, or an auto scaling coverage that hasn\u2019t [&hellip;]","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/","og_site_name":"Future News 24","article_published_time":"2026-06-18T23:31:00+00:00","article_modified_time":"2026-06-19T00:59:24+00:00","og_image":[{"url":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","type":"","width":"","height":""},{"url":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch","datePublished":"2026-06-18T23:31:00+00:00","dateModified":"2026-06-19T00:59:24+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/"},"wordCount":2807,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#primaryimage"},"thumbnailUrl":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","keywords":["CloudWatch","dashboard","debug","detailed","generative","inference","Insights","metrics","Monitor","SageMaker"],"articleSection":["Data Science &amp; MLOps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/","name":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#primaryimage"},"thumbnailUrl":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","datePublished":"2026-06-18T23:31:00+00:00","dateModified":"2026-06-19T00:59:24+00:00","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#primaryimage","url":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png","contentUrl":"https:\/\/d2908q01vomqb2.cloudfront.net\/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59\/2026\/06\/18\/ml-21272.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/18\/monitor-and-debug-generative-ai-inference-with-sagemaker-detailed-metrics-and-insights-dashboard-on-cloudwatch\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1182","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1182"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1182\/revisions"}],"predecessor-version":[{"id":1183,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1182\/revisions\/1183"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1184"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1182"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1182"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1182"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}