We use cookies on our website to give you the most relevant experience by remembering your preferences and repeat visits. By clicking “Accept”, you consent to the use of ALL the cookies. .
Cookie settingsACCEPT
NecessaryAlways Active
Necessary cookies are absolutely essential for the website to function properly. This category only includes cookies that ensures basic functionalities and security features of the website. These cookies do not store any personal information.
- Cookie
__cf_bm
- Duration
1 hour
- Description
This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.
- Cookie
_pxvid
- Duration
1 year
- Description
PerimeterX sets this cookie to detect fraud and bot activity.
- Cookie
_px3
- Duration
6 minutes
- Description
This cookie is set by the Bloomberg to protect the site from BOT attacks.
- Cookie
CookieLawInfoConsent
- Duration
1 year
- Description
CookieYes sets this cookie to record the default button state of the corresponding category and the status of CCPA. It works only in coordination with the primary cookie.
- Cookie
cookielawinfo-checkbox-necessary
- Duration
11 months
- Description
This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
- Cookie
cookielawinfo-checkbox-others
- Duration
1 year
- Description
Set by the GDPR Cookie Consent plugin, this cookie stores user consent for cookies in the category "Others".
- Cookie
cookielawinfo-checkbox-non-necessary
- Duration
11 months
- Description
This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Non Necessary".
- Cookie
cookielawinfo-checkbox-analytics
- Duration
1 year
- Description
Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Analytics" category.
- Cookie
cookielawinfo-checkbox-performance
- Duration
1 year
- Description
Set by the GDPR Cookie Consent plugin, this cookie stores the user consent for cookies in the category "Performance".
- Cookie
cookielawinfo-checkbox-uncategorized
- Duration
1 year
- Description
The cookie is set by the GDPR Cookie Consent plugin to record the user consent for cookies in the category "Uncategorized".
- Cookie
cookielawinfo-checkbox-functional
- Duration
1 year
- Description
The GDPR Cookie Consent plugin sets the cookie to record the user consent for the cookies in the category "Functional".
- Cookie
cookielawinfo-checkbox-advertisement
- Duration
1 year
- Description
Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Advertisement" category.
- Cookie
wpEmojiSettingsSupports
- Duration
session
- Description
WordPress sets this cookie when a user interacts with emojis on a WordPress site. It helps determine if the user's browser can display emojis properly.
- Cookie
VISITOR_PRIVACY_METADATA
- Duration
6 months
- Description
YouTube sets this cookie to store the user's cookie consent state for the current domain.
- Cookie
viewed_cookie_policy
- Duration
11 months
- Description
The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.
- Cookie
PHPSESSID
Duration
Description
This cookie is native to PHP applications. The cookie is used to store and identify a users' unique session ID for the purpose of managing user session on the website. The cookie is a session cookies and is deleted when all the browser windows are closed.
- Cookie
__cfduid
- Duration
4 weeks
- Description
The cookie is set by CloudFare. The cookie is used to identify individual clients behind a shared IP address d apply security settings on a per-client basis. It doesnot correspond to any user ID in the web application and does not store any personally identifiable information.
Functional
Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.
- Cookie
yt-remote-connected-devices
- Duration
never
- Description
YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.
- Cookie
ytidb::LAST_RESULT_ENTRY_KEY
- Duration
never
- Description
The cookie ytidb::LAST_RESULT_ENTRY_KEY is used by YouTube to store the last search result entry that was clicked by the user. This information is used to improve the user experience by providing more relevant search results in the future.
- Cookie
yt-remote-device-id
- Duration
never
- Description
YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.
- Cookie
yt-remote-session-name
- Duration
session
- Description
The yt-remote-session-name cookie is used by YouTube to store the user's video player preferences using embedded YouTube video.
- Cookie
yt-remote-fast-check-period
- Duration
session
- Description
The yt-remote-fast-check-period cookie is used by YouTube to store the user's video player preferences for embedded YouTube videos.
- Cookie
yt-remote-session-app
- Duration
session
- Description
The yt-remote-session-app cookie is used by YouTube to store user preferences and information about the interface of the embedded YouTube video player.
- Cookie
yt-remote-cast-available
- Duration
session
- Description
The yt-remote-cast-available cookie is used to store the user's preferences regarding whether casting is available on their YouTube video player.
- Cookie
yt-remote-cast-installed
- Duration
session
- Description
The yt-remote-cast-installed cookie is used to store the user's video player preferences using embedded YouTube video.
- Cookie
na_id
- Duration
1 year
- Description
This cookie is set by Addthis.com to enable sharing of links on social media platforms like Facebook and Twitter
- Cookie
vc
- Duration
1 year
- Description
This cookie is set by addthis.com on sites that allow sharing on social media.
- Cookie
__atuvc
- Duration
1 year
- Description
This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.
- Cookie
__atuvs
- Duration
30 minutes
- Description
This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.
- Cookie
ouid
- Duration
1 year
- Description
The cookie is set by Addthis which enables the content of the website to be shared across different networking and social sharing websites.
Analytics
Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.
- Cookie
_ga_*
- Duration
1 year 1 month 4 days
- Description
Google Analytics sets this cookie to store and count page views.
- Cookie
_ga
- Duration
2 years
- Description
This cookie is installed by Google Analytics. The cookie is used to calculate visitor, session, camapign data and keep track of site usage for the site's analytics report. The cookies store information anonymously and assigns a randoly generated number to identify unique visitors.
- Cookie
sbjs_migrations
- Duration
session
- Description
Sourcebuster sets this cookie to identify the source of a visit and stores user action information in cookies. This analytical and behavioural cookie is used to enhance the visitor experience on the website.
- Cookie
sbjs_current_add
- Duration
session
Description
Cookie
sbjs_first_add
- Duration
session
Description
Cookie
sbjs_current
- Duration
session
Description
Cookie
sbjs_first
- Duration
session
Description
Cookie
sbjs_udata
- Duration
session
Description
Cookie
sbjs_session
- Duration
1 hour
Description
Cookie
tk_or
- Duration
1 year 1 month 4 days
- Description
JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.
- Cookie
tk_r3d
- Duration
3 days
- Description
JetPack installs this cookie to collect internal metrics for user activity and improve user experience.
- Cookie
tk_lr
- Duration
1 year
- Description
JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.
- Cookie
tk_ai
- Duration
1 year
- Description
JetPack sets this cookie to store a randomly-generated anonymous ID used only within the admin area and for general analytics tracking.
- Cookie
tk_tc
- Duration
session
- Description
JetPack sets this cookie to record details on how users use the website.
- Cookie
_gat_gtag_UA_5784146_31
- Duration
1 minute
- Description
Google Used to distinguish users.
- Cookie
GPS
- Duration
30 minutes
- Description
This cookie is set by Youtube and registers a unique ID for tracking users based on their geographical location
- Cookie
__gads
- Duration
2 years
- Description
This cookie is set by Google and stored under the name dounleclick.com. This cookie is used to track how many times users see a particular advert which helps in measuring the success of the campaign and calculate the revenue generated by the campaign. These cookies can only be read from the domain that it is set on so it will not track any data while browsing through another sites.
- Cookie
uvc
- Duration
1 year
- Description
The cookie is set by addthis.com to determine the usage of Addthis.com service.
- Cookie
ad-id
- Duration
7 months
- Description
Provided by amazon-adsystem.com for tracking user actions on other websites to provide targeted content
- Cookie
_gat_gtag_UA_116563943_1
- Duration
1 minute
- Description
Google uses this cookie to distinguish users.
- Cookie
_gid
- Duration
1 day
- Description
This cookie is installed by Google Analytics. The cookie is used to store information of how visitors use a website and helps in creating an analytics report of how the wbsite is doing. The data collected including the number visitors, the source where they have come from, and the pages viisted in an anonymous form.
Performance
Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.
- Cookie
YSC
Duration
Description
This cookies is set by Youtube and is used to track the views of embedded videos.
- Cookie
_gat
- Duration
1 minute
- Description
This cookies is installed by Google Universal Analytics to throttle the request rate to limit the colllection of data on high traffic sites.
Advertisement
Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.
- Cookie
COMPASS
- Duration
1 hour
- Description
The COMPASS cookie is used by Yahoo to deliver targeted advertising based on user's online behavior.
- Cookie
NID
- Duration
5 months
- Description
This cookie is used to a profile based on user's interest and display personalized ads to the users.
- Cookie
__Secure-YNID
- Duration
6 months
- Description
Google cookie used to protect user security and prevent fraud, especially during the login process.
- Cookie
__Secure-ROLLOUT_TOKEN
- Duration
6 months
- Description
YouTube sets this cookie to manage feature rollout and experimentation. It helps Google control which new features or interface changes are shown to users as part of testing and staged rollouts, ensuring consistent experience for a given user during an experiment.
- Cookie
yt.innertube::nextId
- Duration
never
- Description
YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.
- Cookie
yt.innertube::requests
- Duration
never
- Description
YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.
- Cookie
VISITOR_INFO1_LIVE
- Duration
5 months
- Description
This cookie is set by Youtube. Used to track the information of the embedded YouTube videos on a website.
- Cookie
TapAd_TS
- Duration
1 month
- Description
The cookie is set by Tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising.
- Cookie
TapAd_DID
- Duration
1 month
- Description
The cookie is set by tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising
- Cookie
personalization_id
- Duration
2 years
- Description
This cookie is set by twitter.com. It is used integrate the sharing features of this social media. It also stores information about how the user uses the website for tracking and targeting.
- Cookie
uid
- Duration
1 year
- Description
This cookie is used to measure the number and behavior of the visitors to the website anonymously. The data includes the number of visits, average duration of the visit on the website, pages visited, etc. for the purpose of better understanding user preferences for targeted advertisments.
- Cookie
loc
- Duration
1 year
- Description
This cookie is set by Addthis. This is a geolocation cookie to understand where the users sharing the information are located.
- Cookie
IDE
- Duration
2 years
- Description
Used by Google DoubleClick and stores information about how the user uses the website and any other advertisement before visiting the website. This is used to present users with ads that are relevant to them according to the user profile.
- Cookie
di2
- Duration
1 year
- Description
This cookie is set by addthis.com on sites that allows sharing on social media. The cookie is used to track user behavior anonymously to generate usage trends to improve relevance to their services and advertising.
Others
Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.
- Cookie
pxcts
- Duration
session
- Description
Description is currently not available.
- Cookie
_pxttld
- Duration
session
- Description
Description is currently not available.
- Cookie
SGPBShowingLimitationDomain77659
- Duration
2 days
- Description
Description is currently not available.
- Cookie
__Secure-YEC
- Duration
past
- Description
YouTube sets this cookie to stores the user's video player preferences using embedded YouTube video
- Cookie
S
- Duration
1 hour
- Description
Used by Yahoo to provide ads, content or analytics.
- Cookie
test_cookie
- Duration
11 months
- Description
This cookie is set by doubleclick.net. The purpose of the cookie is to determine if the users' browser supports cookies.
- Cookie
sc_at
- Duration
1 year
- Description
Snapchat sets this cookie for showing relevant advertising based on the user’s movement.
- Cookie
TapAd_3WAY_SYNCS
- Duration
1 month
- Description
TapAd sets this cookie for data synchronization with advertising networks.
- Cookie
_pin_unauth
- Duration
1 year
- Description
Pinterest set this cookie to group actions for users who cannot be identified.
- Cookie
sc_anonymous_id
- Duration
9 years
- Description
Soundcloud sets this cookie to enable visitors to embed content or files on the website.
- Cookie
um
- Duration
1 year
- Description
Set by addthis.com.(Purpose not known)
- Cookie
DCRP_Categories
- Duration
4 weeks
- Description
Description is currently not available.
- Cookie
vuid
- Duration
2 years
- Description
Vimeo installs this cookie to collect tracking information by setting a unique ID to embed videos on the website.
- Cookie
X-AB
- Duration
1 day
- Description
Adobe Analytics sets this cookie in context with multi-variate testing. This is a tool used to combine or change content on the website. This allows the website to find the best variation or edition of the site.
- Cookie
YTC
- Duration
10 minutes
- Description
YouTube sets the YTC cookie to manage the embed and viewing of videos on the website.
- Cookie
sp_t
- Duration
1 month
- Description
The sp_t cookie is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.
- Cookie
sp_landing
- Duration
1 day
- Description
The sp_landing is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.
- Cookie
__asc
- Duration
30 minutes
- Description
Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.
- Cookie
__auc
- Duration
1 year
- Description
Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.
- Cookie
AWSESS
Duration
Description
Awin sets this to ensure the same kind of advertisement is not shown to the user.
- Cookie
nevercache-b39818
- Duration
session
- Description
Description is currently not available.
REJECTSave My PreferencesACCEPT
Powered by
NewsHub](/content/site-root.html)
[Premium Content](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Premium Content"/index.html)
[Read our exclusive articles](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Read our exclusive articles"/index.html)
[Facebook](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Facebook"/index.html)
[Instagram](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Instagram"/index.html)
[X](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "X"/index.html)
Search
NewsHub](/content/site-root.html)
NewsHub](/content/site-root.html)
Search
[Home](/content/ ""/index.html)[Technology](/content/category/technology/ "View all posts in Technology"/index.html)[AI Shorts](/content/category/technology/ai-shorts/ "View all posts in AI Shorts"/index.html)Safely Deploying ML Models to Production: Four Controlled Strategies (A/B, Canary, Interleaved,...
- Technology
- AI Shorts
- Artificial Intelligence
- Applications
- Editors Pick
- Machine Learning
- Staff
- Tech News
- Tutorials
Add as a preferred\ \ source on Google
Deploying a new machine learning model to production is one of the most critical stages of the ML lifecycle. Even if a model performs well on validation and test datasets, directly replacing the existing production model can be risky. Offline evaluation rarely captures the full complexity of real-world environments—data distributions may shift, user behavior can change, and system constraints in production may differ from those in controlled experiments.
As a result, a model that appears superior during development might still degrade performance or negatively impact user experience once deployed. To mitigate these risks, ML teams adopt controlled rollout strategies that allow them to evaluate new models under real production conditions while minimizing potential disruptions.
In this article, we explore four widely used strategies—A/B testing, Canary testing, Interleaved testing, and Shadow testing—that help organizations safely deploy and validate new machine learning models in production environments.
A/B Testing
A/B testing is one of the most widely used strategies for safely introducing a new machine learning model in production. In this approach, incoming traffic is split between two versions of a system: the existing legacy model (control) and the candidate model (variation). The distribution is typically non-uniform to limit risk—for example, 90% of requests may continue to be served by the legacy model, while only 10% are routed to the candidate model.
By exposing both models to real-world traffic, teams can compare downstream performance metrics such as click-through rate, conversions, engagement, or revenue. This controlled experiment allows organizations to evaluate whether the candidate model genuinely improves outcomes before gradually increasing its traffic share or fully replacing the legacy model.
Canary Testing
Canary testing is a controlled rollout strategy where a new model is first deployed to a small subset of users before being gradually released to the entire user base. The name comes from an old mining practice where miners carried canary birds into coal mines to detect toxic gases—the birds would react first, warning miners of danger. Similarly, in machine learning deployments, the candidate model is initially exposed to a limited group of users while the majority continue to be served by the legacy model.
Unlike A/B testing, which randomly splits traffic across all users, canary testing targets a specific subset and progressively increases exposure if performance metrics indicate success. This gradual rollout helps teams detect issues early and roll back quickly if necessary, reducing the risk of widespread impact.
Interleaved Testing
Interleaved testing evaluates multiple models by mixing their outputs within the same response shown to users. Instead of routing an entire request to either the legacy or candidate model, the system combines predictions from both models in real time. For example, in a recommendation system, some items in the recommendation list may come from the legacy model, while others are generated by the candidate model.
The system then logs downstream engagement signals—such as click-through rate, watch time, or negative feedback—for each recommendation. Because both models are evaluated within the same user interaction, interleaved testing allows teams to compare performance more directly and efficiently while minimizing biases caused by differences in user groups or traffic distribution.
Shadow Testing
Shadow testing, also known as shadow deployment or dark launch, allows teams to evaluate a new machine learning model in a real production environment without affecting the user experience. In this approach, the candidate model runs in parallel with the legacy model and receives the same live requests as the production system. However, only the legacy model’s predictions are returned to users, while the candidate model’s outputs are simply logged for analysis.
This setup helps teams assess how the new model behaves under real-world traffic and infrastructure conditions, which are often difficult to replicate in offline experiments. Shadow testing provides a low-risk way to benchmark the candidate model against the legacy model, although it cannot capture true user engagement metrics—such as clicks, watch time, or conversions—since its predictions are never shown to users.
Simulating ML Model Deployment Strategies
Setting Up
Before simulating any strategy, we need two things: a way to represent incoming requests, and a stand-in for each model.
Each model is simply a function that takes a request and returns a score — a number that loosely represents how good that model’s recommendation is. The legacy model’s score is capped at 0.35, while the candidate model’s is capped at 0.55, making the candidate intentionally better so we can verify that each strategy actually detects the improvement.
make_requests() generates 200 requests spread across 40 users, which gives us enough traffic to see meaningful differences between strategies while keeping the simulation lightweight.
Copy CodeCopiedUse a different Browser
A/B Testing
ab_route() is the core of this strategy — for every incoming request, it draws a random number and routes to the candidate model only if that number falls below 0.10, otherwise the request goes to legacy. This gives the candidate roughly 10% of traffic.
We then collect the prediction scores from each model separately and compute the average at the end. In a real system, these scores would be replaced by actual engagement metrics like click-through rate or watch time — here the score just stands in for “how good was this recommendation.”
Copy CodeCopiedUse a different Browser
Canary Testing
The key function here is get_canary_users(), which uses an MD5 hash to deterministically assign users to the canary group. The important word is deterministic — sorting users by their hash means the same users always end up in the canary group across runs, which mirrors how real canary deployments work where a specific user consistently sees the same model.
We then simulate three phases by simply expanding the fraction of canary users — 5%, 20%, and 50%. For each request, routing is decided by whether the user belongs to the canary group, not by a random coin flip like in A/B testing. This is the fundamental difference between the two strategies: A/B testing splits by request, canary testing splits by user.
Copy CodeCopiedUse a different Browser
Interleaved Testing
Both models run on every request, and interleave() merges their outputs by alternating items — one from legacy, one from candidate, one from legacy, and so on. Each item is tagged with its source model, so when a user clicks something, we know exactly which model to credit.
The small random.uniform(-0.05, 0.05) noise added to each item’s score simulates the natural variation you’d see in real recommendations — two items from the same model won’t have identical quality.
At the end, we compute CTR separately for each model’s items. Because both models competed on the same requests against the same users at the same time, there is no confounding factor — any difference in CTR is purely down to model quality. This is what makes interleaved testing the most statistically clean comparison of the four strategies.
Copy CodeCopiedUse a different Browser
Shadow Testing
Both models run on every request, but the loop makes a clear distinction — live_pred is what the user gets, shadow_pred goes straight into the log and nothing more. The candidate’s output is never returned, never shown, never acted on. The log list is the entire point of shadow testing. In a real system this would be written to a database or a data warehouse, and engineers would later query it to compare latency distributions, output patterns, or score distributions against the legacy model — all without a single user being affected.
Copy CodeCopiedUse a different Browser
Check out the FULL Notebook Here. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
View Arham Islam's Linkedin profile
Arham Islam
I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.
- Arham Islam
Stochastic Gradient Descent (SGD’s) Frequency Bias and How Adam Fixes It
- Arham Islam
Understanding LLM Distillation Techniques
- Arham Islam
Why Gradient Descent Zigzags and How Momentum Fixes It
- Arham Islam
- Arham Islam
What is Tokenization Drift and How to Fix It?
- Arham Islam
The LoRA Assumption That Breaks in Production
- Arham Islam
RAG Without Vectors: How PageIndex Retrieves by Reasoning
- Arham Islam
- Arham Islam
- Arham Islam
How Knowledge Distillation Compresses Ensemble Intelligence into a Single Deployable AI Model
- Arham Islam
Five AI Compute Architectures Every Engineer Should Know: CPUs, GPUs, TPUs, NPUs, and LPUs Compared
- Arham Islam
Sigmoid vs ReLU Activation Functions: The Inference Cost of Losing Geometric Context
- Arham Islam
Paged Attention in Large Language Models LLMs
- Arham Islam
How BM25 and RAG Retrieve Information Differently?
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
Getting Started with OpenClaw and Connecting It with WhatsApp
- Arham Islam
The Statistical Cost of Zero Padding in Convolutional Neural Networks (CNNs)
- Arham Islam
- Arham Islam
Understanding the Layers of AI Observability in the Age of LLMs
- Arham Islam
Implementing Softmax From Scratch: Avoiding the Numerical Stability Trap
- Arham Islam
AI Interview Series #5: Prompt Caching
- Arham Islam
AI Interview Series #4: Explain KV Caching
- Arham Islam
- Arham Islam
5 AI Model Architectures Every AI Engineer Should Know
- Arham Islam
Kernel Principal Component Analysis (PCA): Explained with an Example
- Arham Islam
AI Interview Series #4: Transformers vs Mixture of Experts (MoE)
- Arham Islam
AI Interview Series #3: Explain Federated Learning
- Arham Islam
Focal Loss vs Binary Cross-Entropy: A Practical Guide for Imbalanced Classification
- Arham Islam
- Arham Islam
How to Reduce Cost and Latency of Your RAG Application Using Semantic LLM Caching
- Arham Islam
AI Interview Series #1: Explain Some LLM Text Generation Strategies Used in LLMs
- Arham Islam
How to Build Supervised AI Models When You Don’t Have Annotated Data
- Arham Islam
- Arham Islam
- Arham Islam
5 Common LLM Parameters Explained with Examples
- Arham Islam
- Arham Islam
A Guide for Effective Context Engineering for AI Agents
- Arham Islam
How to Evaluate Your RAG Pipeline with Synthetic Data?
- Arham Islam
5 Most Popular Agentic AI Design Patterns Every AI Engineer Should Know
- Arham Islam
Building a Human Handoff Interface for AI-Powered Insurance Agent Using Parlant and Streamlit
- Arham Islam
Agentic Design Methodology: How to Build Reliable and Human-Like AI Agents using Parlant
- Arham Islam
Ensuring AI Safety in Production: A Developer’s Guide to OpenAI’s Moderation and Safety Checks
- Arham Islam
- Arham Islam
How to Create Reliable Conversational AI Agents Using Parlant?
- Arham Islam
Understanding the Universal Tool Calling Protocol (UTCP)
- Arham Islam
Top 5 No-Code Tools for AI Engineers/Developers
- Arham Islam
Implementing OAuth 2.1 for MCP Servers with Scalekit: A Step-by-Step Coding Tutorial
- Arham Islam
- Arham Islam
How to Implement the LLM Arena-as-a-Judge Approach to Evaluate Large Language Model Outputs
- Arham Islam
JSON Prompting for LLMs: A Practical Guide with Python Coding Examples
- Arham Islam
Creating Dashboards Using Vizro MCP: Vizro is an Open-Source Python Toolkit by McKinsey
- Arham Islam
How to Test an OpenAI Model Against Single-Turn Adversarial Attacks Using deepteam
- Arham Islam
Using RouteLLM to Optimize LLM Usage
- Arham Islam
A Developer’s Guide to OpenAI’s GPT-5 Model Capabilities
- Arham Islam
Tutorial: Exploring SHAP-IQ Visualizations
- Arham Islam
- Arham Islam
Implementing Self-Refine Technique Using Large Language Models LLMs
- Arham Islam
Creating a Knowledge Graph Using an LLM
- Arham Islam
o1 Style Thinking with Chain-of-Thought Reasoning using Mirascope
- Arham Islam
Getting Started with Mirascope: Removing Semantic Duplicates using an LLM
- Arham Islam
Tracing OpenAI Agent Responses using MLFlow
- Arham Islam
Getting Started with Agent Communication Protocol (ACP): Build a Weather Agent with Python
- Arham Islam
Getting started with Gemini Command Line Interface (CLI)
- Arham Islam
Getting Started with MLFlow for LLM Evaluation
- Arham Islam
- Arham Islam
Teaching Mistral Agents to Say No: Content Moderation from Prompt to Response
- Arham Islam
- Arham Islam
- Arham Islam
How to Create Smart Multi-Agent Workflows Using the Mistral Agents API’s Handoffs Feature
- Arham Islam
How to Enable Function Calling in Mistral Agents Using the Standard JSON Schema Format
- Arham Islam
Hands-On Guide: Getting started with Mistral Agents API
- Arham Islam
Guide to Using the Desktop Commander MCP Server
- Arham Islam
Step-by-Step Guide to Creating Synthetic Data Using the Synthetic Data Vault (SDV)
- Arham Islam
Step-by-Step Guide to Create an AI agent with Google ADK
- Arham Islam
Implementing an LLM Agent with Tool Access Using MCP-Use
- Arham Islam
Implementing an AgentQL Model Context Protocol (MCP) Server
- Arham Islam
Implementing An Airbnb and Excel MCP Server
- Arham Islam
How to Create a Custom Model Context Protocol (MCP) Client Using Gemini
- Arham Islam
Implementing Persistent Memory Using a Local Knowledge Graph in Claude Desktop
- Arham Islam
Step by Step Guide on How to Convert a FastAPI App into an MCP Server
- Arham Islam
Integrating Figma with Cursor IDE Using an MCP Server to Build a Web Login Page
- Arham Islam
- Arham Islam
40+ Cool AI Tools You Should Check Out (Oct 2024)
- Arham Islam
- Arham Islam
Meta AI Researchers Open-Source Pearl: A Production-Ready Reinforcement Learning AI Agent Library
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
This AI Paper Unveils DiffEnc: Advancing Diffusion Models for Enhanced Generative Performance
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
50+ New Cutting-Edge Artificial Intelligence AI Tools (November 2023)
- Arham Islam
List of Artificial Intelligence AI Advancements by Non-Profit Researchers
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
From 2D to 3D: Enhancing Text-to-3D Generation Consistency with Aligned Geometric Priors
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
- Arham Islam
LLMs and Data Analysis: How AI is Making Sense of Big Data for Business Insights
- Arham Islam
Role of Data Contracts in Data Pipeline
- Arham Islam
40+ AI Tools For Video Creation and Editing in 2023
- Arham Islam
Artificial Intelligence (AI) and Web3: How are they Connected?
- Arham Islam
- Arham Islam
52 AI Tools For Sales Professionals (2023)
- Arham Islam
Use of Analog Computers in Artificial Intelligence (AI)
- Arham Islam
- Arham Islam
Multimodal Language Models: The Future of Artificial Intelligence (AI)
- Arham Islam
Top 50+ AI Coding Assistant Tools in 2023
- Arham Islam
List of Groundbreaking and Open-Source Conversational AI Models in the Language Domain
- Arham Islam
Application of Large Language Models in Biotechnology and Pharmaceutical Research
- Arham Islam
Top 50+ AI Tools for Marketers 2023
- Arham Islam
The Groundbreaking Influence of Generative AI in the Automotive Industry
- Arham Islam
What is Field Programmable Gate Array (FPGA): FPGA vs. GPU for Artificial Intelligence (AI)
- Arham Islam
10 Use Cases of ChatGPT in Marketing for 2023
- Arham Islam
Meet AIAgent: A Web-based AutomateGPT that Needs No API Keys and is Powered by GPT4
- Arham Islam
Exploring the Benefits and Drawbacks of Integrating ChatGPT into Healthcare
- Arham Islam
How To Use Third-Party Plugins In ChatGPT? 80+ Plugins Just Added by ChatGPT For Public
- Arham Islam
- Arham Islam
7 AI Tools that Transform Anything into Interactive Chatbots
- Arham Islam
Meet Window AI: A New Way To Use Your Own AI Models On The Web – Including Local Ones
- Arham Islam
12 Creative Ways Developers Can Use Chat GPT-4
- Arham Islam
A History of Generative AI: From GAN to GPT-4
- Arham Islam
Roadmap of Becoming a Prompt Engineer (2023)
- Arham Islam
What is ChatGPT? Technology Behind ChatGPT
- Arham Islam
- Arham Islam
- Arham Islam
5 GANs Concepts You Should Know About in 2023
- Arham Islam
What are Transformers? Concept and Applications Explained
- Arham Islam
Best Practices For Machine Learning Model Monitoring
- Arham Islam
- Arham Islam
Bad Data Engineering Practices And How To Avoid Them
- Arham Islam
What is Multimodal Learning? Some Applications
- Arham Islam
High-Performance Computing (HPC) And Artificial Intelligence (AI)
- Arham Islam
What is Dataops (Data Operations)? Difference between DataOps and DevOps
- Arham Islam
How Do DALL·E 2, Stable Diffusion, and Midjourney Work?
- Arham Islam
What is AIOps (Artificial Intelligence for IT Operations)?AIOps Use Cases
- Arham Islam
- Arham Islam
AI Hardware Accelerators For Machine Learning And Deep Learning | How To Choose One
- Arham Islam
- Arham Islam
Understanding The Artificial Intelligence (AI) Bill of Rights From The White House
- Arham Islam
Top Real World Applications of Reinforcement Learning in 2022
RELATED ARTICLES MORE FROM AUTHOR
[How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)
[Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order](/content/2026/06/13/anthropic-disables-claude-fable-5-and-mythos-5-after-us-government-order/ "Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order"/index.html)
[Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)
[A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)
[Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard](/content/2026/06/12/google-releases-gemini-sql2-gemini-3-1-pro-text-to-sql-scores-80-04-on-bird-single-model-leaderboard/ "Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard"/index.html)
[Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)
[How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access,...](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)
Sana Hassan-June 13, 20260
In this tutorial, we implement a QwenPaw workflow that provides a practical environment for building and testing an agent-powered assistant. We install and initialize...
Asif Razzaq-June 13, 20260
shutdown followed a US government export control directive citing national security authorities. All other Anthropic models, including Opus 4.8, remain available.
[Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench...](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)
Asif Razzaq-June 12, 20260
Moonshot AI has open-sourced Kimi K2.7-Code under a Modified MIT license. It is a coding-focused, agentic model built on Kimi K2.6, with a 256K context window and roughly 30% lower reasoning-token usage. Moonshot reports gains over K2.6 on six benchmarks, including +21.8% on Kimi Code Bench v2. The model is available via the Kimi API and Kimi Code.
[A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph,...](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)
Sana Hassan-June 12, 20260
We build an end-to-end spatial graph learning pipeline using city2graph. We collect urban POI and street network data from OpenStreetMap, with a synthetic fallback for reliability. We engineer spatial features, construct several proximity graph families, and compare how each represents the same urban environment. We then build heterogeneous and homogeneous graphs, convert them to PyTorch Geometric, and train a GraphSAGE model to predict POI categories from spatial structure.
Asif Razzaq-June 12, 20260
We look at Gemini-SQL2, the text-to-SQL capability Google Research announced on June 12, 2026. Powered by Gemini 3.1 Pro, it posted 80.04% execution accuracy on the BIRD single-model leaderboard. We explain what the score measures, how the leaderboard stacks up, and what Google has not yet disclosed. We also cover use cases and a schema-grounded implementation pattern.
[Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6...](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)
Asif Razzaq-June 12, 20260
Moonshot AI's Kimi Work is a local desktop agent for macOS and Windows. It runs a 300-sub-agent swarm, drives your logged-in browser via WebBridge, and schedules background jobs.
[Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order...](/content/2026/06/12/zyphra-release-zamba2-vl-hybrid-mamba2-transformer-vision-language-models-that-cut-time-to-first-token-by-about-an-order-of-magnitude/ "Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude"/index.html)
Asif Razzaq-June 12, 20260
Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about an order of magnitude.
[A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical...](/content/2026/06/12/a-coding-implementation-on-monai-for-end-to-end-3d-spleen-segmentation-using-unet-on-medical-ct-volumes/ "A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical CT Volumes"/index.html)
Sana Hassan-June 12, 20260
In this tutorial, we build an end-to-end 3D medical image segmentation pipeline using MONAI to segment the spleen on the Medical Segmentation Decathlon Task09...
[Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For...](/content/2026/06/11/perplexity-moves-deep-research-into-computer-routing-research-subtasks-across-20-frontier-models-for-reports-decks-and-dashboards/ "Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For Reports, Decks, And Dashboards"/index.html)
Michal Sutter-June 11, 20260
Deep Research now lives inside Perplexity Computer, breaking hard questions into subtasks and routing across 20+ frontier models.
[xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and...](/content/2026/06/11/xai-ships-grok-build-plugin-marketplace-with-mongodb-vercel-sentry-chrome-devtools-cloudflare-and-superpowers-plugins-at-launch/ "xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and Superpowers Plugins at Launch"/index.html)
Michal Sutter-June 11, 20260
Grok Build's in-terminal marketplace bundles skills, agents, hooks, and MCP servers, with commit-SHA verification on every remote plugin.
© Copyright Reserved @2025 Marktechpost AI Media Inc
Toggle photo metadata visibilityToggle photo comments visibility
Loading Comments...
Write a Comment...
Email (Required)Name (Required)Website