We use cookies on our website to give you the most relevant experience by remembering your preferences and repeat visits. By clicking “Accept”, you consent to the use of ALL the cookies. .

Cookie settingsACCEPT

NecessaryAlways Active

Necessary cookies are absolutely essential for the website to function properly. This category only includes cookies that ensures basic functionalities and security features of the website. These cookies do not store any personal information.

- Cookie

\_\_cf\_bm

- Duration

1 hour

- Description

This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.

- Cookie

\_pxvid

- Duration

1 year

- Description

PerimeterX sets this cookie to detect fraud and bot activity.

- Cookie

\_px3

- Duration

6 minutes

- Description

This cookie is set by the Bloomberg to protect the site from BOT attacks.

- Cookie

CookieLawInfoConsent

- Duration

1 year

- Description

CookieYes sets this cookie to record the default button state of the corresponding category and the status of CCPA. It works only in coordination with the primary cookie.

- Cookie

cookielawinfo-checkbox-necessary

- Duration

11 months

- Description

This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".

- Cookie

cookielawinfo-checkbox-others

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie stores user consent for cookies in the category "Others".

- Cookie

cookielawinfo-checkbox-non-necessary

- Duration

11 months

- Description

This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Non Necessary".

- Cookie

cookielawinfo-checkbox-analytics

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Analytics" category.

- Cookie

cookielawinfo-checkbox-performance

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie stores the user consent for cookies in the category "Performance".

- Cookie

cookielawinfo-checkbox-uncategorized

- Duration

1 year

- Description

The cookie is set by the GDPR Cookie Consent plugin to record the user consent for cookies in the category "Uncategorized".

- Cookie

cookielawinfo-checkbox-functional

- Duration

1 year

- Description

The GDPR Cookie Consent plugin sets the cookie to record the user consent for the cookies in the category "Functional".

- Cookie

cookielawinfo-checkbox-advertisement

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Advertisement" category.

- Cookie

wpEmojiSettingsSupports

- Duration

session

- Description

WordPress sets this cookie when a user interacts with emojis on a WordPress site. It helps determine if the user's browser can display emojis properly.

- Cookie

VISITOR\_PRIVACY\_METADATA

- Duration

6 months

- Description

YouTube sets this cookie to store the user's cookie consent state for the current domain.

- Cookie

viewed\_cookie\_policy

- Duration

11 months

- Description

The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.

- Cookie

PHPSESSID

- Duration

- Description

This cookie is native to PHP applications. The cookie is used to store and identify a users' unique session ID for the purpose of managing user session on the website. The cookie is a session cookies and is deleted when all the browser windows are closed.

- Cookie

\_\_cfduid

- Duration

4 weeks

- Description

The cookie is set by CloudFare. The cookie is used to identify individual clients behind a shared IP address d apply security settings on a per-client basis. It doesnot correspond to any user ID in the web application and does not store any personally identifiable information.

Functional

Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.

- Cookie

yt-remote-connected-devices

- Duration

never

- Description

YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.

- Cookie

ytidb::LAST\_RESULT\_ENTRY\_KEY

- Duration

never

- Description

The cookie ytidb::LAST\_RESULT\_ENTRY\_KEY is used by YouTube to store the last search result entry that was clicked by the user. This information is used to improve the user experience by providing more relevant search results in the future.

- Cookie

yt-remote-device-id

- Duration

never

- Description

YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.

- Cookie

yt-remote-session-name

- Duration

session

- Description

The yt-remote-session-name cookie is used by YouTube to store the user's video player preferences using embedded YouTube video.

- Cookie

yt-remote-fast-check-period

- Duration

session

- Description

The yt-remote-fast-check-period cookie is used by YouTube to store the user's video player preferences for embedded YouTube videos.

- Cookie

yt-remote-session-app

- Duration

session

- Description

The yt-remote-session-app cookie is used by YouTube to store user preferences and information about the interface of the embedded YouTube video player.

- Cookie

yt-remote-cast-available

- Duration

session

- Description

The yt-remote-cast-available cookie is used to store the user's preferences regarding whether casting is available on their YouTube video player.

- Cookie

yt-remote-cast-installed

- Duration

session

- Description

The yt-remote-cast-installed cookie is used to store the user's video player preferences using embedded YouTube video.

- Cookie

na\_id

- Duration

1 year

- Description

This cookie is set by Addthis.com to enable sharing of links on social media platforms like Facebook and Twitter

- Cookie

vc

- Duration

1 year

- Description

This cookie is set by addthis.com on sites that allow sharing on social media.

- Cookie

\_\_atuvc

- Duration

1 year

- Description

This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.

- Cookie

\_\_atuvs

- Duration

30 minutes

- Description

This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.

- Cookie

ouid

- Duration

1 year

- Description

The cookie is set by Addthis which enables the content of the website to be shared across different networking and social sharing websites.

Analytics

Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.

- Cookie

\_ga\_\*

- Duration

1 year 1 month 4 days

- Description

Google Analytics sets this cookie to store and count page views.

- Cookie

\_ga

- Duration

2 years

- Description

This cookie is installed by Google Analytics. The cookie is used to calculate visitor, session, camapign data and keep track of site usage for the site's analytics report. The cookies store information anonymously and assigns a randoly generated number to identify unique visitors.

- Cookie

sbjs\_migrations

- Duration

session

- Description

Sourcebuster sets this cookie to identify the source of a visit and stores user action information in cookies. This analytical and behavioural cookie is used to enhance the visitor experience on the website.

- Cookie

sbjs\_current\_add

- Duration

session

- Description

- Cookie

sbjs\_first\_add

- Duration

session

- Description

- Cookie

sbjs\_current

- Duration

session

- Description

- Cookie

sbjs\_first

- Duration

session

- Description

- Cookie

sbjs\_udata

- Duration

session

- Description

- Cookie

sbjs\_session

- Duration

1 hour

- Description

- Cookie

tk\_or

- Duration

1 year 1 month 4 days

- Description

JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.

- Cookie

tk\_r3d

- Duration

3 days

- Description

JetPack installs this cookie to collect internal metrics for user activity and improve user experience.

- Cookie

tk\_lr

- Duration

1 year

- Description

JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.

- Cookie

tk\_ai

- Duration

1 year

- Description

JetPack sets this cookie to store a randomly-generated anonymous ID used only within the admin area and for general analytics tracking.

- Cookie

tk\_tc

- Duration

session

- Description

JetPack sets this cookie to record details on how users use the website.

- Cookie

\_gat\_gtag\_UA\_5784146\_31

- Duration

1 minute

- Description

Google Used to distinguish users.

- Cookie

GPS

- Duration

30 minutes

- Description

This cookie is set by Youtube and registers a unique ID for tracking users based on their geographical location

- Cookie

\_\_gads

- Duration

2 years

- Description

This cookie is set by Google and stored under the name dounleclick.com. This cookie is used to track how many times users see a particular advert which helps in measuring the success of the campaign and calculate the revenue generated by the campaign. These cookies can only be read from the domain that it is set on so it will not track any data while browsing through another sites.

- Cookie

uvc

- Duration

1 year

- Description

The cookie is set by addthis.com to determine the usage of Addthis.com service.

- Cookie

ad-id

- Duration

7 months

- Description

Provided by amazon-adsystem.com for tracking user actions on other websites to provide targeted content

- Cookie

\_gat\_gtag\_UA\_116563943\_1

- Duration

1 minute

- Description

Google uses this cookie to distinguish users.

- Cookie

\_gid

- Duration

1 day

- Description

This cookie is installed by Google Analytics. The cookie is used to store information of how visitors use a website and helps in creating an analytics report of how the wbsite is doing. The data collected including the number visitors, the source where they have come from, and the pages viisted in an anonymous form.

Performance

Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.

- Cookie

YSC

- Duration

- Description

This cookies is set by Youtube and is used to track the views of embedded videos.

- Cookie

\_gat

- Duration

1 minute

- Description

This cookies is installed by Google Universal Analytics to throttle the request rate to limit the colllection of data on high traffic sites.

Advertisement

Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.

- Cookie

COMPASS

- Duration

1 hour

- Description

The COMPASS cookie is used by Yahoo to deliver targeted advertising based on user's online behavior.

- Cookie

NID

- Duration

5 months

- Description

This cookie is used to a profile based on user's interest and display personalized ads to the users.

- Cookie

\_\_Secure-YNID

- Duration

6 months

- Description

Google cookie used to protect user security and prevent fraud, especially during the login process.

- Cookie

\_\_Secure-ROLLOUT\_TOKEN

- Duration

6 months

- Description

YouTube sets this cookie to manage feature rollout and experimentation. It helps Google control which new features or interface changes are shown to users as part of testing and staged rollouts, ensuring consistent experience for a given user during an experiment.

- Cookie

yt.innertube::nextId

- Duration

never

- Description

YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.

- Cookie

yt.innertube::requests

- Duration

never

- Description

YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.

- Cookie

VISITOR\_INFO1\_LIVE

- Duration

5 months

- Description

This cookie is set by Youtube. Used to track the information of the embedded YouTube videos on a website.

- Cookie

TapAd\_TS

- Duration

1 month

- Description

The cookie is set by Tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising.

- Cookie

TapAd\_DID

- Duration

1 month

- Description

The cookie is set by tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising

- Cookie

personalization\_id

- Duration

2 years

- Description

This cookie is set by twitter.com. It is used integrate the sharing features of this social media. It also stores information about how the user uses the website for tracking and targeting.

- Cookie

uid

- Duration

1 year

- Description

This cookie is used to measure the number and behavior of the visitors to the website anonymously. The data includes the number of visits, average duration of the visit on the website, pages visited, etc. for the purpose of better understanding user preferences for targeted advertisments.

- Cookie

loc

- Duration

1 year

- Description

This cookie is set by Addthis. This is a geolocation cookie to understand where the users sharing the information are located.

- Cookie

IDE

- Duration

2 years

- Description

Used by Google DoubleClick and stores information about how the user uses the website and any other advertisement before visiting the website. This is used to present users with ads that are relevant to them according to the user profile.

- Cookie

di2

- Duration

1 year

- Description

This cookie is set by addthis.com on sites that allows sharing on social media. The cookie is used to track user behavior anonymously to generate usage trends to improve relevance to their services and advertising.

Others

Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.

- Cookie

pxcts

- Duration

session

- Description

Description is currently not available.

- Cookie

\_pxttld

- Duration

session

- Description

Description is currently not available.

- Cookie

SGPBShowingLimitationDomain77659

- Duration

2 days

- Description

Description is currently not available.

- Cookie

\_\_Secure-YEC

- Duration

past

- Description

YouTube sets this cookie to stores the user's video player preferences using embedded YouTube video

- Cookie

S

- Duration

1 hour

- Description

Used by Yahoo to provide ads, content or analytics.

- Cookie

test\_cookie

- Duration

11 months

- Description

This cookie is set by doubleclick.net. The purpose of the cookie is to determine if the users' browser supports cookies.

- Cookie

sc\_at

- Duration

1 year

- Description

Snapchat sets this cookie for showing relevant advertising based on the user’s movement.

- Cookie

TapAd\_3WAY\_SYNCS

- Duration

1 month

- Description

TapAd sets this cookie for data synchronization with advertising networks.

- Cookie

\_pin\_unauth

- Duration

1 year

- Description

Pinterest set this cookie to group actions for users who cannot be identified.

- Cookie

sc\_anonymous\_id

- Duration

9 years

- Description

Soundcloud sets this cookie to enable visitors to embed content or files on the website.

- Cookie

um

- Duration

1 year

- Description

Set by addthis.com.(Purpose not known)

- Cookie

DCRP\_Categories

- Duration

4 weeks

- Description

Description is currently not available.

- Cookie

vuid

- Duration

2 years

- Description

Vimeo installs this cookie to collect tracking information by setting a unique ID to embed videos on the website.

- Cookie

X-AB

- Duration

1 day

- Description

Adobe Analytics sets this cookie in context with multi-variate testing. This is a tool used to combine or change content on the website. This allows the website to find the best variation or edition of the site.

- Cookie

YTC

- Duration

10 minutes

- Description

YouTube sets the YTC cookie to manage the embed and viewing of videos on the website.

- Cookie

sp\_t

- Duration

1 month

- Description

The sp\_t cookie is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.

- Cookie

sp\_landing

- Duration

1 day

- Description

The sp\_landing is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.

- Cookie

\_\_asc

- Duration

30 minutes

- Description

Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.

- Cookie

\_\_auc

- Duration

1 year

- Description

Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.

- Cookie

AWSESS

- Duration

- Description

Awin sets this to ensure the same kind of advertisement is not shown to the user.

- Cookie

nevercache-b39818

- Duration

session

- Description

Description is currently not available.

REJECTSave My PreferencesACCEPT

Powered by

NewsHub](/content/site-root.html)

[Premium Content](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Premium Content"/index.html)

[Read our exclusive articles](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Read our exclusive articles"/index.html)

[Facebook](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Facebook"/index.html)

[Instagram](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "Instagram"/index.html)

[X](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/# "X"/index.html)

[Discord](https://pxl.to/ivxz41s "Discord")[Linkedin](https://www.linkedin.com/company/marktechpost/?viewAsMember=true "Linkedin")[Reddit](https://www.reddit.com/r/machinelearningnews/ "Reddit")[X](https://twitter.com/Marktechpost "X")

- [Home](/content/site-root.html)
- [Open Source/Weights](/content/category/technology/open-source/index.html)
- [AI Agents](/content/category/editors-pick/ai-agents/index.html)
- [Tutorials](/content/category/tutorials/index.html)
- [Voice AI](/content/category/technology/artificial-intelligence/voice-ai/index.html)
- [Robotics](/content/category/robotics/index.html)
- [Newsletter](https://www.aidevsignals.com/)
- [→ Partner with Us](https://forms.gle/CY1eqZzuWFQBp7dH9)

Search

NewsHub](/content/site-root.html)

NewsHub](/content/site-root.html)

Search

[Home](/content/ ""/index.html)[Technology](/content/category/technology/ "View all posts in Technology"/index.html)[AI Shorts](/content/category/technology/ai-shorts/ "View all posts in AI Shorts"/index.html)Safely Deploying ML Models to Production: Four Controlled Strategies (A/B, Canary, Interleaved,...

[tinyfish.aiOpen Source\\
\\
Big **Set**\\
\\
Describe your ideal dataset in plain English, and BigSet builds it.\\
\\
dataset.build()auto·refresh\\
\\
✓\\
\\
✓\\
\\
✓\\
\\
✓\\
\\
Explore on GitHub→](https://pxllnk.co/cuv4rk8)

- [Technology](/content/category/technology/index.html)
- [AI Shorts](/content/category/technology/ai-shorts/index.html)
- [Artificial Intelligence](/content/category/technology/artificial-intelligence/index.html)
- [Applications](/content/category/technology/artificial-intelligence/applications/index.html)
- [Editors Pick](/content/category/editors-pick/index.html)
- [Machine Learning](/content/category/technology/artificial-intelligence/machine-learning/index.html)
- [Staff](/content/category/editors-pick/staff/index.html)
- [Tech News](/content/category/tech-news/index.html)
- [Tutorials](/content/category/tutorials/index.html)

[Add as a preferred\\
\\
source on Google](https://www.google.com/preferences/source?q=https://www.marktechpost.com/)

Deploying a new machine learning model to production is one of the most critical stages of the ML lifecycle. Even if a model performs well on validation and test datasets, directly replacing the existing production model can be risky. Offline evaluation rarely captures the full complexity of real-world environments—data distributions may shift, user behavior can change, and system constraints in production may differ from those in controlled experiments.

As a result, a model that appears superior during development might still degrade performance or negatively impact user experience once deployed. To mitigate these risks, ML teams adopt controlled rollout strategies that allow them to evaluate new models under real production conditions while minimizing potential disruptions.

In this article, we explore four widely used strategies—A/B testing, Canary testing, Interleaved testing, and Shadow testing—that help organizations safely deploy and validate new machine learning models in production environments.

## **A/B Testing**

**A/B testing** is one of the most widely used strategies for safely introducing a new machine learning model in production. In this approach, incoming traffic is split between two versions of a system: the existing **legacy model** (control) and the **candidate model** (variation). The distribution is typically non-uniform to limit risk—for example, 90% of requests may continue to be served by the legacy model, while only 10% are routed to the candidate model.

By exposing both models to real-world traffic, teams can compare downstream performance metrics such as click-through rate, conversions, engagement, or revenue. This controlled experiment allows organizations to evaluate whether the candidate model genuinely improves outcomes before gradually increasing its traffic share or fully replacing the legacy model.

## **Canary Testing**

**Canary testing** is a controlled rollout strategy where a new model is first deployed to a small subset of users before being gradually released to the entire user base. The name comes from an old mining practice where miners carried canary birds into coal mines to detect toxic gases—the birds would react first, warning miners of danger. Similarly, in machine learning deployments, the **candidate model** is initially exposed to a limited group of users while the majority continue to be served by the **legacy model**.

Unlike A/B testing, which randomly splits traffic across all users, canary testing targets a specific subset and progressively increases exposure if performance metrics indicate success. This gradual rollout helps teams detect issues early and roll back quickly if necessary, reducing the risk of widespread impact.

## **Interleaved Testing**

**Interleaved testing** evaluates multiple models by mixing their outputs within the same response shown to users. Instead of routing an entire request to either the legacy or candidate model, the system combines predictions from both models in real time. For example, in a recommendation system, some items in the recommendation list may come from the **legacy model**, while others are generated by the **candidate model**.

The system then logs downstream engagement signals—such as click-through rate, watch time, or negative feedback—for each recommendation. Because both models are evaluated within the same user interaction, interleaved testing allows teams to compare performance more directly and efficiently while minimizing biases caused by differences in user groups or traffic distribution.

## **Shadow Testing**

**Shadow testing**, also known as **shadow deployment** or **dark launch**, allows teams to evaluate a new machine learning model in a real production environment without affecting the user experience. In this approach, the **candidate model** runs in parallel with the **legacy model** and receives the same live requests as the production system. However, only the legacy model’s predictions are returned to users, while the candidate model’s outputs are simply logged for analysis.

This setup helps teams assess how the new model behaves under real-world traffic and infrastructure conditions, which are often difficult to replicate in offline experiments. Shadow testing provides a low-risk way to benchmark the candidate model against the legacy model, although it cannot capture true user engagement metrics—such as clicks, watch time, or conversions—since its predictions are never shown to users.

## **Simulating ML Model Deployment Strategies**

### **Setting Up**

Before simulating any strategy, we need two things: a way to represent incoming requests, and a stand-in for each model.

Each model is simply a function that takes a request and returns a score — a number that loosely represents how good that model’s recommendation is. The legacy model’s score is capped at 0.35, while the candidate model’s is capped at 0.55, making the candidate intentionally better so we can verify that each strategy actually detects the improvement.

make\_requests() generates 200 requests spread across 40 users, which gives us enough traffic to see meaningful differences between strategies while keeping the simulation lightweight.

Copy CodeCopiedUse a different Browser

```php

```

### **A/B Testing**

ab\_route() is the core of this strategy — for every incoming request, it draws a random number and routes to the candidate model only if that number falls below 0.10, otherwise the request goes to legacy. This gives the candidate roughly 10% of traffic.

We then collect the prediction scores from each model separately and compute the average at the end. In a real system, these scores would be replaced by actual engagement metrics like click-through rate or watch time — here the score just stands in for “how good was this recommendation.”

Copy CodeCopiedUse a different Browser

```php

```

##

### **Canary Testing**

The key function here is get\_canary\_users(), which uses an MD5 hash to deterministically assign users to the canary group. The important word is deterministic — sorting users by their hash means the same users always end up in the canary group across runs, which mirrors how real canary deployments work where a specific user consistently sees the same model.

We then simulate three phases by simply expanding the fraction of canary users — 5%, 20%, and 50%. For each request, routing is decided by whether the user belongs to the canary group, not by a random coin flip like in A/B testing. This is the fundamental difference between the two strategies: A/B testing splits by request, canary testing splits by user.

Copy CodeCopiedUse a different Browser

```php

```

##

### **Interleaved Testing**

Both models run on every request, and interleave() merges their outputs by alternating items — one from legacy, one from candidate, one from legacy, and so on. Each item is tagged with its source model, so when a user clicks something, we know exactly which model to credit.

The small random.uniform(-0.05, 0.05) noise added to each item’s score simulates the natural variation you’d see in real recommendations — two items from the same model won’t have identical quality.

At the end, we compute CTR separately for each model’s items. Because both models competed on the same requests against the same users at the same time, there is no confounding factor — any difference in CTR is purely down to model quality. This is what makes interleaved testing the most statistically clean comparison of the four strategies.

Copy CodeCopiedUse a different Browser

```php

```

### **Shadow Testing**

Both models run on every request, but the loop makes a clear distinction — live\_pred is what the user gets, shadow\_pred goes straight into the log and nothing more. The candidate’s output is never returned, never shown, never acted on. The log list is the entire point of shadow testing. In a real system this would be written to a database or a data warehouse, and engineers would later query it to compare latency distributions, output patterns, or score distributions against the legacy model — all without a single user being affected.

Copy CodeCopiedUse a different Browser

```php

```

* * *

Check out the **[FULL Notebook Here](https://github.com/Marktechpost/AI-Tutorial-Codes-Included/blob/main/Data%20Science/ML_Deployment_Methods.ipynb).** Also, feel free to follow us on **[Twitter](https://x.com/intent/follow?screen_name=marktechpost)** and don’t forget to join our **[120k+ ML SubReddit](https://www.reddit.com/r/machinelearningnews/)** and Subscribe to **[our Newsletter](https://www.aidevsignals.com/)**. Wait! are you on telegram? **[now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)**

[View Arham Islam's Linkedin profile](https://www.linkedin.com/in/mohd-arham-islam/)

##### [Arham Islam](/content/author/arhamislam/index.html)

[\+ postsBio](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/#/index.html)

I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.

- Arham Islam

[Stochastic Gradient Descent (SGD’s) Frequency Bias and How Adam Fixes It](/content/2026/05/18/stochastic-gradient-descent-sgds-frequency-bias-and-how-adam-fixes-it/index.html)

- Arham Islam

[Understanding LLM Distillation Techniques](/content/2026/05/11/understanding-llm-distillation-techniques/index.html)

- Arham Islam

[Why Gradient Descent Zigzags and How Momentum Fixes It](/content/2026/05/05/why-gradient-descent-zigzags-and-how-momentum-fixes-it/index.html)

- Arham Islam

[A Developer’s Guide to Systematic Prompting: Mastering Negative Constraints, Structured JSON Outputs, and Multi-Hypothesis Verbalized Sampling](/content/2026/05/03/a-developers-guide-to-systematic-prompting-mastering-negative-constraints-structured-json-outputs-and-multi-hypothesis-verbalized-sampling/index.html)

- Arham Islam

[What is Tokenization Drift and How to Fix It?](/content/2026/05/03/what-is-tokenization-drift-and-how-to-fix-it/index.html)

- Arham Islam

[The LoRA Assumption That Breaks in Production](/content/2026/04/26/the-lora-assumption-that-breaks-in-production/index.html)

- Arham Islam

[RAG Without Vectors: How PageIndex Retrieves by Reasoning](/content/2026/04/25/rag-without-vectors-how-pageindex-retrieves-by-reasoning/index.html)

- Arham Islam

[How TabPFN Leverages In-Context Learning to Achieve Superior Accuracy on Tabular Datasets Compared to Random Forest and CatBoost](/content/2026/04/19/how-tabpfn-leverages-in-context-learning-to-achieve-superior-accuracy-on-tabular-datasets-compared-to-random-forest-and-catboost/index.html)

- Arham Islam

[A Technical Deep Dive into the Essential Stages of Modern Large Language Model Training, Alignment, and Deployment](/content/2026/04/15/a-technical-deep-dive-into-the-essential-stages-of-modern-large-language-model-training-alignment-and-deployment/index.html)

- Arham Islam

[How Knowledge Distillation Compresses Ensemble Intelligence into a Single Deployable AI Model](/content/2026/04/11/how-knowledge-distillation-compresses-ensemble-intelligence-into-a-single-deployable-ai-model/index.html)

- Arham Islam

[Five AI Compute Architectures Every Engineer Should Know: CPUs, GPUs, TPUs, NPUs, and LPUs Compared](/content/2026/04/09/five-ai-compute-architectures-every-engineer-should-know-cpus-gpus-tpus-npus-and-lpus-compared/index.html)

- Arham Islam

[Sigmoid vs ReLU Activation Functions: The Inference Cost of Losing Geometric Context](/content/2026/04/09/sigmoid-vs-relu-activation-functions-the-inference-cost-of-losing-geometric-context/index.html)

- Arham Islam

[Paged Attention in Large Language Models LLMs](/content/2026/03/24/paged-attention-in-large-language-models-llms/index.html)

- Arham Islam

[How BM25 and RAG Retrieve Information Differently?](/content/2026/03/22/how-bm25-and-rag-retrieve-information-differently/index.html)

- Arham Islam

[Model Context Protocol (MCP) vs. AI Agent Skills: A Deep Dive into Structured Tools and Behavioral Guidance for LLMs](/content/2026/03/13/model-context-protocol-mcp-vs-ai-agent-skills-a-deep-dive-into-structured-tools-and-behavioral-guidance-for-llms/index.html)

- Arham Islam

[Beyond Accuracy: Quantifying the Production Fragility Caused by Excessive, Redundant, and Low-Signal Features in Regression](/content/2026/03/08/beyond-accuracy-quantifying-the-production-fragility-caused-by-excessive-redundant-and-low-signal-features-in-regression/index.html)

- Arham Islam

[RAG vs. Context Stuffing: Why selective retrieval is more efficient and reliable than dumping all data into the prompt](/content/2026/02/24/rag-vs-context-stuffing-why-selective-retrieval-is-more-efficient-and-reliable-than-dumping-all-data-into-the-prompt/index.html)

- Arham Islam

[Getting Started with OpenClaw and Connecting It with WhatsApp](/content/2026/02/14/getting-started-with-openclaw-and-connecting-it-with-whatsapp/index.html)

- Arham Islam

[The Statistical Cost of Zero Padding in Convolutional Neural Networks (CNNs)](/content/2026/02/02/the-statistical-cost-of-zero-padding-in-convolutional-neural-networks-cnns/index.html)

- Arham Islam

[What are Context Graphs?](/content/2026/01/20/what-are-context-graphs/index.html)

- Arham Islam

[Understanding the Layers of AI Observability in the Age of LLMs](/content/2026/01/13/understanding-the-layers-of-ai-observability-in-the-age-of-llms/index.html)

- Arham Islam

[Implementing Softmax From Scratch: Avoiding the Numerical Stability Trap](/content/2026/01/06/implementing-softmax-from-scratch-avoiding-the-numerical-stability-trap/index.html)

- Arham Islam

[AI Interview Series #5: Prompt Caching](/content/2026/01/04/ai-interview-series-5-prompt-caching/index.html)

- Arham Islam

[AI Interview Series #4: Explain KV Caching](/content/2025/12/21/ai-interview-series-4-explain-kv-caching/index.html)

- Arham Islam

[Google Introduces T5Gemma 2: Encoder Decoder Models with Multimodal Inputs via SigLIP and 128K Context](/content/2025/12/19/google-introduces-t5gemma-2-encoder-decoder-models-with-multimodal-inputs-via-siglip-and-128k-context/index.html)

- Arham Islam

[5 AI Model Architectures Every AI Engineer Should Know](/content/2025/12/12/5-ai-model-architectures-every-ai-engineer-should-know/index.html)

- Arham Islam

[Kernel Principal Component Analysis (PCA): Explained with an Example](/content/2025/12/05/kernel-principal-component-analysis-pca-explained-with-an-example/index.html)

- Arham Islam

[AI Interview Series #4: Transformers vs Mixture of Experts (MoE)](/content/2025/12/03/ai-interview-series-4-transformers-vs-mixture-of-experts-moe/index.html)

- Arham Islam

[AI Interview Series #3: Explain Federated Learning](/content/2025/11/23/ai-interview-series-3-explain-federated-learning/index.html)

- Arham Islam

[Focal Loss vs Binary Cross-Entropy: A Practical Guide for Imbalanced Classification](/content/2025/11/17/focal-loss-vs-binary-cross-entropy-a-practical-guide-for-imbalanced-classification/index.html)

- Arham Islam

[AI Interview Series #2: Explain Some of the Common Model Context Protocol (MCP) Security Vulnerabilities](/content/2025/11/16/ai-interview-series-2-explain-some-of-the-common-model-context-protocol-mcp-security-vulnerabilities/index.html)

- Arham Islam

[How to Reduce Cost and Latency of Your RAG Application Using Semantic LLM Caching](/content/2025/11/11/how-to-reduce-cost-and-latency-of-your-rag-application-using-semantic-llm-caching/index.html)

- Arham Islam

[AI Interview Series #1: Explain Some LLM Text Generation Strategies Used in LLMs](/content/2025/11/09/ai-interview-series-1-explain-some-llm-text-generation-strategies-used-in-llms/index.html)

- Arham Islam

[How to Build Supervised AI Models When You Don’t Have Annotated Data](/content/2025/11/03/how-to-build-supervised-ai-models-when-you-dont-have-annotated-data/index.html)

- Arham Islam

[How to Create AI-ready APIs?](/content/2025/11/02/how-to-create-ai-ready-apis/index.html)

- Arham Islam

[Meet Pyversity Library: How to Improve Retrieval Systems by Diversifying the Results Using Pyversity?](/content/2025/10/27/meet-pyversity-library-how-to-improve-retrieval-systems-by-diversifying-the-results-using-pyversity/index.html)

- Arham Islam

[5 Common LLM Parameters Explained with Examples](/content/2025/10/26/5-common-llm-parameters-explained-with-examples/index.html)

- Arham Islam

[Meet LangChain’s DeepAgents Library and a Practical Example to See How DeepAgents Actually Work in Action](/content/2025/10/20/meet-langchains-deepagents-library-and-a-practical-example-to-see-how-deepagents-actually-work-in-action/index.html)

- Arham Islam

[A Guide for Effective Context Engineering for AI Agents](/content/2025/10/20/a-guide-for-effective-context-engineering-for-ai-agents/index.html)

- Arham Islam

[How to Evaluate Your RAG Pipeline with Synthetic Data?](/content/2025/10/13/how-to-evaluate-your-rag-pipeline-with-synthetic-data/index.html)

- Arham Islam

[5 Most Popular Agentic AI Design Patterns Every AI Engineer Should Know](/content/2025/10/12/5-most-popular-agentic-ai-design-patterns-every-ai-engineer-should-know/index.html)

- Arham Islam

[Building a Human Handoff Interface for AI-Powered Insurance Agent Using Parlant and Streamlit](/content/2025/10/06/building-a-human-handoff-interface-for-ai-powered-insurance-agent-using-parlant-and-streamlit/index.html)

- Arham Islam

[Agentic Design Methodology: How to Build Reliable and Human-Like AI Agents using Parlant](/content/2025/10/05/agentic-design-methodology-how-to-build-reliable-and-human-like-ai-agents-using-parlant/index.html)

- Arham Islam

[Ensuring AI Safety in Production: A Developer’s Guide to OpenAI’s Moderation and Safety Checks](/content/2025/09/28/ensuring-ai-safety-in-production-a-developers-guide-to-openais-moderation-and-safety-checks/index.html)

- Arham Islam

[What is Asyncio? Getting Started with Asynchronous Python and Using Asyncio in an AI Application with an LLM](/content/2025/09/27/what-is-asyncio-getting-started-with-asynchronous-python-and-using-asyncio-in-an-ai-application-with-an-llm/index.html)

- Arham Islam

[How to Create Reliable Conversational AI Agents Using Parlant?](/content/2025/09/22/how-to-create-reliable-conversational-ai-agents-using-parlant/index.html)

- Arham Islam

[Understanding the Universal Tool Calling Protocol (UTCP)](/content/2025/09/21/understanding-the-universal-tool-calling-protocol-utcp/index.html)

- Arham Islam

[Top 5 No-Code Tools for AI Engineers/Developers](/content/2025/09/14/top-5-no-code-tools-for-ai-engineers-developers/index.html)

- Arham Islam

[Implementing OAuth 2.1 for MCP Servers with Scalekit: A Step-by-Step Coding Tutorial](/content/2025/09/01/implementing-oauth-2-1-for-mcp-servers-with-scalekit-a-step-by-step-coding-tutorial/index.html)

- Arham Islam

[Understanding OAuth 2.1 for MCP (Model Context Protocol) Servers: Discovery, Authorization, and Access Phases](/content/2025/08/31/understanding-oauth-2-1-for-mcp-model-context-protocol-servers-discovery-authorization-and-access-phases/index.html)

- Arham Islam

[How to Implement the LLM Arena-as-a-Judge Approach to Evaluate Large Language Model Outputs](/content/2025/08/25/how-to-implement-the-llm-arena-as-a-judge-approach-to-evaluate-large-language-model-outputs/index.html)

- Arham Islam

[JSON Prompting for LLMs: A Practical Guide with Python Coding Examples](/content/2025/08/23/json-prompting-for-llms-a-practical-guide-with-python-coding-examples/index.html)

- Arham Islam

[Creating Dashboards Using Vizro MCP: Vizro is an Open-Source Python Toolkit by McKinsey](/content/2025/08/18/creating-dashboards-using-vizro-mcp-vizro-is-an-open-source-python-toolkit-by-mckinsey/index.html)

- Arham Islam

[How to Test an OpenAI Model Against Single-Turn Adversarial Attacks Using deepteam](/content/2025/08/17/how-to-test-an-openai-model-against-single-turn-adversarial-attacks-using-deepteam/index.html)

- Arham Islam

[Using RouteLLM to Optimize LLM Usage](/content/2025/08/10/using-routellm-to-optimize-llm-usage/index.html)

- Arham Islam

[A Developer’s Guide to OpenAI’s GPT-5 Model Capabilities](/content/2025/08/08/a-developers-guide-to-openais-gpt-5-model-capabilities/index.html)

- Arham Islam

[Tutorial: Exploring SHAP-IQ Visualizations](/content/2025/08/03/tutorial-exploring-shap-iq-visualizations/index.html)

- Arham Islam

[How to Use the SHAP-IQ Package to Uncover and Visualize Feature Interactions in Machine Learning Models Using Shapley Interaction Indices (SII)](/content/2025/08/02/how-to-use-the-shap-iq-package-to-uncover-and-visualize-feature-interactions-in-machine-learning-models-using-shapley-interaction-indices-sii/index.html)

- Arham Islam

[Implementing Self-Refine Technique Using Large Language Models LLMs](/content/2025/07/29/implementing-self-refine-technique-using-large-language-models-llms/index.html)

- Arham Islam

[Creating a Knowledge Graph Using an LLM](/content/2025/07/28/creating-a-knowledge-graph-using-an-llm/index.html)

- Arham Islam

[o1 Style Thinking with Chain-of-Thought Reasoning using Mirascope](/content/2025/07/18/o1-style-thinking-with-chain-of-thought-reasoning-using-mirascope/index.html)

- Arham Islam

[Getting Started with Mirascope: Removing Semantic Duplicates using an LLM](/content/2025/07/16/getting-started-with-mirascope-removing-semantic-duplicates-using-an-llm/index.html)

- Arham Islam

[Tracing OpenAI Agent Responses using MLFlow](/content/2025/07/14/tracing-openai-agent-responses-using-mlflow/index.html)

- Arham Islam

[Getting Started with Agent Communication Protocol (ACP): Build a Weather Agent with Python](/content/2025/07/06/getting-started-with-agent-communication-protocol-acp-build-a-weather-agent-with-python/index.html)

- Arham Islam

[Getting started with Gemini Command Line Interface (CLI)](/content/2025/06/28/getting-started-with-gemini-command-line-interface-cli/index.html)

- Arham Islam

[Getting Started with MLFlow for LLM Evaluation](/content/2025/06/27/getting-started-with-mlflow-for-llm-evaluation/index.html)

- Arham Islam

[Getting Started with Microsoft’s Presidio: A Step-by-Step Guide to Detecting and Anonymizing Personally Identifiable Information PII in Text](/content/2025/06/24/getting-started-with-microsofts-presidio-a-step-by-step-guide-to-detecting-and-anonymizing-personally-identifiable-information-pii-in-text/index.html)

- Arham Islam

[Teaching Mistral Agents to Say No: Content Moderation from Prompt to Response](/content/2025/06/23/teaching-mistral-agents-to-say-no-content-moderation-from-prompt-to-response/index.html)

- Arham Islam

[Building an A2A-Compliant Random Number Agent: A Step-by-Step Guide to Implementing the Low-Level Executor Pattern with Python](/content/2025/06/21/building-an-a2a-compliant-random-number-agent-a-step-by-step-guide-to-implementing-the-low-level-executor-pattern-with-python/index.html)

- Arham Islam

[How to Use python-A2A to Create and Connect Financial Agents with Google’s Agent-to-Agent (A2A) Protocol](/content/2025/06/16/how-to-use-python-a2a-to-create-and-connect-financial-agents-with-googles-agent-to-agent-a2a-protocol/index.html)

- Arham Islam

[How to Create Smart Multi-Agent Workflows Using the Mistral Agents API’s Handoffs Feature](/content/2025/06/09/how-to-create-smart-multi-agent-workflows-using-the-mistral-agents-apis-handoffs-feature/index.html)

- Arham Islam

[How to Enable Function Calling in Mistral Agents Using the Standard JSON Schema Format](/content/2025/06/08/how-to-enable-function-calling-in-mistral-agents-using-the-standard-json-schema-format/index.html)

- Arham Islam

[Hands-On Guide: Getting started with Mistral Agents API](/content/2025/06/03/hands-on-guide-getting-started-with-mistral-agents-api/index.html)

- Arham Islam

[Guide to Using the Desktop Commander MCP Server](/content/2025/06/01/guide-to-using-the-desktop-commander-mcp-server/index.html)

- Arham Islam

[Step-by-Step Guide to Creating Synthetic Data Using the Synthetic Data Vault (SDV)](/content/2025/05/25/step-by-step-guide-to-creating-synthetic-data-using-the-synthetic-data-vault-sdv/index.html)

- Arham Islam

[Step-by-Step Guide to Create an AI agent with Google ADK](/content/2025/05/20/step-by-step-guide-to-create-an-ai-agent-with-google-adk/index.html)

- Arham Islam

[Implementing an LLM Agent with Tool Access Using MCP-Use](/content/2025/05/13/implementing-an-llm-agent-with-tool-access-using-mcp-use/index.html)

- Arham Islam

[Implementing an AgentQL Model Context Protocol (MCP) Server](/content/2025/05/06/implementing-an-agentql-model-context-protocol-mcp-server/index.html)

- Arham Islam

[Implementing An Airbnb and Excel MCP Server](/content/2025/05/02/implementing-an-airbnb-and-excel-mcp-server/index.html)

- Arham Islam

[How to Create a Custom Model Context Protocol (MCP) Client Using Gemini](/content/2025/04/29/how-to-create-a-custom-model-context-protocol-mcp-client-using-gemini/index.html)

- Arham Islam

[Implementing Persistent Memory Using a Local Knowledge Graph in Claude Desktop](/content/2025/04/26/implementing-persistent-memory-using-a-local-knowledge-graph-in-claude-desktop/index.html)

- Arham Islam

[Step by Step Guide on How to Convert a FastAPI App into an MCP Server](/content/2025/04/19/step-by-step-guide-on-how-to-convert-a-fastapi-app-into-an-mcp-server/index.html)

- Arham Islam

[Integrating Figma with Cursor IDE Using an MCP Server to Build a Web Login Page](/content/2025/04/17/integrating-figma-with-cursor-ide-using-an-mcp-server-to-build-a-web-login-page/index.html)

- Arham Islam

[Code Implementation to Building a Model Context Protocol (MCP) Server and Connecting It with Claude Desktop](/content/2025/04/13/code-implementation-to-building-a-model-context-protocol-mcp-server-and-connecting-it-with-claude-desktop/index.html)

- Arham Islam

[40+ Cool AI Tools You Should Check Out (Oct 2024)](/content/2024/10/11/40-cool-ai-tools-you-should-check-out-december-2023/index.html)

- Arham Islam

[Pinterest Researchers Present an Effective Scalable Algorithm to Improve Diffusion Models Using Reinforcement Learning (RL)](/content/2024/02/11/pinterest-researchers-present-an-effective-scalable-algorithm-to-improve-diffusion-models-using-reinforcement-learning-rl/index.html)

- Arham Islam

[Meta AI Researchers Open-Source Pearl: A Production-Ready Reinforcement Learning AI Agent Library](/content/2023/12/11/meta-ai-researchers-open-source-pearl-a-production-ready-reinforcement-learning-ai-agent-library/index.html)

- Arham Islam

[Researchers from the University of Texas Showcase Predicting Implant-Based Reconstruction Complications Using Machine Learning](/content/2023/12/04/researchers-from-the-university-of-texas-showcase-predicting-implant-based-reconstruction-complications-using-machine-learning/index.html)

- Arham Islam

[UC Berkeley Researchers Propose an Artificial Intelligence Algorithm that Achieves Zero-Shot Acquisition of Goal-Directed Dialogue Agents](/content/2023/11/18/uc-berkeley-researchers-propose-an-artificial-intelligence-algorithm-that-achieves-zero-shot-acquisition-of-goal-directed-dialogue-agents/index.html)

- Arham Islam

[Can Language Models Reason Beyond Words? Exploring Implicit Reasoning in Multi-Layer Hidden States for Complex Tasks](/content/2023/11/15/can-language-models-reason-beyond-words-exploring-implicit-reasoning-in-multi-layer-hidden-states-for-complex-tasks/index.html)

- Arham Islam

[Meta Researchers Introduced VR-NeRF: An Advanced End-to-End AI System for High-Fidelity Capture and Rendering of Walkable Spaces in Virtual Reality](/content/2023/11/12/meta-researchers-introduced-vr-nerf-an-advanced-end-to-end-ai-system-for-high-fidelity-capture-and-rendering-of-walkable-spaces-in-virtual-reality/index.html)

- Arham Islam

[Are You Doing Retrieval-Augmented Generation (RAG) for Biomedicine? Meet MedCPT: A Contrastive Pre-trained Transformer Model for Zero-Shot Biomedical Information Retrieval](/content/2023/11/11/are-you-doing-retrieval-augmented-generation-rag-for-biomedicine-meet-medcpt-a-contrastive-pre-trained-transformer-model-for-zero-shot-biomedical-information-retrieval/index.html)

- Arham Islam

[Intel Researchers Propose a New Artificial Intelligence Approach to Deploy LLMs on CPUs More Efficiently](/content/2023/11/09/intel-researchers-propose-a-new-artificial-intelligence-approach-to-deploy-llms-on-cpus-more-efficiently/index.html)

- Arham Islam

[This AI Paper Unveils DiffEnc: Advancing Diffusion Models for Enhanced Generative Performance](/content/2023/11/07/this-ai-paper-unveils-diffenc-advancing-diffusion-models-for-enhanced-generative-performance/index.html)

- Arham Islam

[A New AI Research from China Introduces GLM-130B: A Bilingual (English and Chinese) Pre-Trained Language Model with 130B Parameters](/content/2023/11/02/a-new-ai-research-from-china-introduces-glm-130b-a-bilingual-english-and-chinese-pre-trained-language-model-with-130b-parameters/index.html)

- Arham Islam

[Unlocking the Secrets of CLIP’s Data Success: Introducing MetaCLIP for Optimized Language-Image Pre-training](/content/2023/10/31/unlocking-the-secrets-of-clips-data-success-introducing-metaclip-for-optimized-language-image-pre-training/index.html)

- Arham Islam

[Researchers from the University of Washington and Princeton Present a Pre-Training Data Detection Dataset WIKIMIA and a New Machine Learning Approach MIN-K% PROB](/content/2023/10/30/researchers-from-the-university-of-washington-and-princeton-present-a-pre-training-data-detection-dataset-wikimia-and-a-new-machine-learning-approach-min-k-prob/index.html)

- Arham Islam

[50+ New Cutting-Edge Artificial Intelligence AI Tools (November 2023)](/content/2023/10/30/50-new-cutting-edge-ai-tools-2023/index.html)

- Arham Islam

[List of Artificial Intelligence AI Advancements by Non-Profit Researchers](/content/2023/10/27/list-of-artificial-intelligence-ai-advancements-by-non-profit-researchers/index.html)

- Arham Islam

[Revolutionizing Language Model Fine-Tuning: Achieving Unprecedented Gains with NEFTune’s Noisy Embeddings](/content/2023/10/24/revolutionizing-language-model-fine-tuning-achieving-unprecedented-gains-with-neftunes-noisy-embeddings/index.html)

- Arham Islam

[A New AI Research from China Proposes 4K4D: A 4D Point Cloud Representation that Supports Hardware Rasterization and Enables Unprecedented Rendering Speed](/content/2023/10/21/a-new-ai-research-from-china-proposes-4k4d-a-4d-point-cloud-representation-that-supports-hardware-rasterization-and-enables-unprecedented-rendering-speed/index.html)

- Arham Islam

[This AI Paper Proposes ‘MotionDirector’: An Artificial Intelligence Approach to Customize Video Motion and Appearance](/content/2023/10/17/this-ai-paper-proposes-motiondirector-an-artificial-intelligence-approach-to-customize-video-motion-and-appearance/index.html)

- Arham Islam

[From 2D to 3D: Enhancing Text-to-3D Generation Consistency with Aligned Geometric Priors](/content/2023/10/16/from-2d-to-3d-enhancing-text-to-3d-generation-consistency-with-aligned-geometric-priors/index.html)

- Arham Islam

[Google AI Introduces SANPO: A Multi-Attribute Video Dataset for Outdoor Human Egocentric Scene Understanding](/content/2023/10/14/google-ai-introduces-sanpo-a-multi-attribute-video-dataset-for-outdoor-human-egocentric-scene-understanding/index.html)

- Arham Islam

[This AI Research Proposes Kosmos-G: An Artificial Intelligence Model that Performs High-Fidelity Zero-Shot Image Generation from Generalized Vision-Language Input Leveraging the property of Multimodel LLMs](/content/2023/10/11/this-ai-research-proposes-kosmos-g-an-artificial-intelligence-model-that-performs-high-fidelity-zero-shot-image-generation-from-generalized-vision-language-input-leveraging-the-property-of-multimodel/index.html)

- Arham Islam

[Latest Advancements in the Field of Multimodal AI: (ChatGPT + DALLE 3) + (Google BARD + Extensions) and many more….](/content/2023/10/05/latest-advancements-in-the-field-of-multimodal-ai-chatgpt-dalle-3-google-bard-extensions-and-many-more/index.html)

- Arham Islam

[What is Model Merging?](/content/2023/09/27/what-is-model-merging/index.html)

- Arham Islam

[LLMs & Knowledge Graphs](/content/2023/09/19/llms-knowledge-graphs/index.html)

- Arham Islam

[LLMs and Data Analysis: How AI is Making Sense of Big Data for Business Insights](/content/2023/09/11/llms-and-data-analysis-how-ai-is-making-sense-of-big-data-for-business-insights/index.html)

- Arham Islam

[Role of Data Contracts in Data Pipeline](/content/2023/08/26/role-of-data-contracts-in-data-pipeline/index.html)

- Arham Islam

[40+ AI Tools For Video Creation and Editing in 2023](/content/2023/08/22/40-ai-tools-for-video-creation-and-editing-in-2023/index.html)

- Arham Islam

[Artificial Intelligence (AI) and Web3: How are they Connected?](/content/2023/08/07/artificial-intelligence-ai-and-web3-how-are-they-connected/index.html)

- Arham Islam

[LLMs Outperform Reinforcement Learning- Meet SPRING: An Innovative Prompting Framework for LLMs Designed to Enable in-Context Chain-of-Thought Planning and Reasoning](/content/2023/08/01/llms-outperform-reinforcement-learning-meet-spring-an-innovative-prompting-framework-for-llms-designed-to-enable-in-context-chain-of-thought-planning-and-reasoning/index.html)

- Arham Islam

[52 AI Tools For Sales Professionals (2023)](/content/2023/07/25/52-ai-tools-for-sales-professionals-2023/index.html)

- Arham Islam

[Use of Analog Computers in Artificial Intelligence (AI)](/content/2023/07/24/use-of-analog-computers-in-artificial-intelligence-ai/index.html)

- Arham Islam

[A New AI Research Presents A Prompt-Centric Approach For Analyzing Large Language Models LLMs Capabilities](/content/2023/07/22/a-new-ai-research-presents-a-prompt-centric-approach-for-analyzing-large-language-models-llms-capabilities/index.html)

- Arham Islam

[Multimodal Language Models: The Future of Artificial Intelligence (AI)](/content/2023/07/19/multimodal-language-models-the-future-of-artificial-intelligence-ai/index.html)

- Arham Islam

[Top 50+ AI Coding Assistant Tools in 2023](/content/2023/07/18/top-50-ai-coding-assistant-tools-in-2023/index.html)

- Arham Islam

[List of Groundbreaking and Open-Source Conversational AI Models in the Language Domain](/content/2023/07/15/list-of-groundbreaking-and-open-source-conversational-ai-models-in-the-language-domain/index.html)

- Arham Islam

[Application of Large Language Models in Biotechnology and Pharmaceutical Research](/content/2023/07/12/application-of-large-language-models-in-biotechnology-and-pharmaceutical-research/index.html)

- Arham Islam

[Top 50+ AI Tools for Marketers 2023](/content/2023/07/10/top-50-ai-tools-for-marketers-2023/index.html)

- Arham Islam

[The Groundbreaking Influence of Generative AI in the Automotive Industry](/content/2023/07/03/the-groundbreaking-influence-of-generative-ai-in-the-automotive-industry/index.html)

- Arham Islam

[What is Field Programmable Gate Array (FPGA): FPGA vs. GPU for Artificial Intelligence (AI)](/content/2023/07/02/what-is-field-programmable-gate-array-fpga-fpga-vs-gpu-for-artificial-intelligence-ai/index.html)

- Arham Islam

[10 Use Cases of ChatGPT in Marketing for 2023](/content/2023/06/30/10-use-cases-of-chatgpt-in-marketing-for-2023/index.html)

- Arham Islam

[Meet AIAgent: A Web-based AutomateGPT that Needs No API Keys and is Powered by GPT4](/content/2023/06/19/meet-aiagent-a-web-based-automategpt-that-needs-no-api-keys-and-is-powered-by-gpt4/index.html)

- Arham Islam

[Exploring the Benefits and Drawbacks of Integrating ChatGPT into Healthcare](/content/2023/06/13/exploring-the-benefits-and-drawbacks-of-integrating-chatgpt-into-healthcare/index.html)

- Arham Islam

[How To Use Third-Party Plugins In ChatGPT? 80+ Plugins Just Added by ChatGPT For Public](/content/2023/05/21/how-to-use-third-party-plugins-in-chatgpt-80-plugins-just-added-by-chatgpt-for-public/index.html)

- Arham Islam

[Google Just Announced “Help Me Write” Feature in Gmail: AI Creates An Email With Just One Line Prompt](/content/2023/05/14/google-just-announced-help-me-write-feature-in-gmail-ai-creates-an-email-with-just-one-line-prompt/index.html)

- Arham Islam

[7 AI Tools that Transform Anything into Interactive Chatbots](/content/2023/04/23/5-ai-tools-that-transform-anything-into-interactive-chatbots/index.html)

- Arham Islam

[Meet Window AI: A New Way To Use Your Own AI Models On The Web – Including Local Ones](/content/2023/04/22/meet-window-ai-a-new-way-to-use-your-own-ai-models-on-the-web-including-local-ones/index.html)

- Arham Islam

[12 Creative Ways Developers Can Use Chat GPT-4](/content/2023/04/11/12-creative-ways-developers-can-use-chat-gpt-4/index.html)

- Arham Islam

[A History of Generative AI: From GAN to GPT-4](/content/2023/03/21/a-history-of-generative-ai-from-gan-to-gpt-4/index.html)

- Arham Islam

[Roadmap of Becoming a Prompt Engineer (2023)](/content/2023/03/12/roadmap-of-becoming-a-prompt-engineer-2023/index.html)

- Arham Islam

[What is ChatGPT? Technology Behind ChatGPT](/content/2023/03/04/what-is-chatgpt-technology-behind-chatgpt/index.html)

- Arham Islam

[Top Large Language Models (LLMs) in 2023 from OpenAI, Google AI, Deepmind, Anthropic, Baidu, Huawei, Meta AI, AI21 Labs, LG AI Research and NVIDIA](/content/2023/02/22/top-large-language-models-llms-in-2023-from-openai-google-ai-deepmind-anthropic-baidu-huawei-meta-ai-ai21-labs-lg-ai-research-and-nvidia/index.html)

- Arham Islam

[A New Prompt Engineering Research Proposes PEZ (Prompts Made Easy): A Gradient Optimizer For Text That Utilizes Continuous Embeddings To Reliably Optimize Hard Prompts](/content/2023/02/10/a-new-prompt-engineering-research-proposes-pez-prompts-made-easy-a-gradient-optimizer-for-text-that-utilizes-continuous-embeddings-to-reliably-optimize-hard-prompts/index.html)

- Arham Islam

[5 GANs Concepts You Should Know About in 2023](/content/2023/02/04/5-gans-concepts-you-should-know-about-in-2023/index.html)

- Arham Islam

[What are Transformers? Concept and Applications Explained](/content/2023/01/24/what-are-transformers-concept-and-applications-explained/index.html)

- Arham Islam

[Best Practices For Machine Learning Model Monitoring](/content/2023/01/22/best-practices-for-machine-learning-model-monitoring/index.html)

- Arham Islam

[Artificial Intelligence (AI) Research Innovations in 2022 from Google, NVIDIA, Salesforce, Meta, Apple, Amazon, and AI2](/content/2023/01/10/artificial-intelligence-ai-research-innovations-in-2022-from-google-nvidia-salesforce-meta-apple-amazon-and-ai2/index.html)

- Arham Islam

[Bad Data Engineering Practices And How To Avoid Them](/content/2022/12/20/bad-data-engineering-practices-and-how-to-avoid-them/index.html)

- Arham Islam

[What is Multimodal Learning? Some Applications](/content/2022/12/12/what-is-multimodal-learning-some-applications/index.html)

- Arham Islam

[High-Performance Computing (HPC) And Artificial Intelligence (AI)](/content/2022/12/04/high-performance-computing-hpc-and-artificial-intelligence-ai/index.html)

- Arham Islam

[What is Dataops (Data Operations)? Difference between DataOps and DevOps](/content/2022/11/27/what-is-dataops-data-operations-difference-between-dataops-and-devops/index.html)

- Arham Islam

[How Do DALL·E 2, Stable Diffusion, and Midjourney Work?](/content/2022/11/14/how-do-dall%c2%b7e-2-stable-diffusion-and-midjourney-work/index.html)

- Arham Islam

[What is AIOps (Artificial Intelligence for IT Operations)?AIOps Use Cases](/content/2022/11/05/what-is-aiops-artificial-intelligence-for-it-operationsaiops-use-cases/index.html)

- Arham Islam

[What is MLOps (Machine Learning Operations)? Why Do You Need MLOps for Machine Learning and Deep Learning Projects?](/content/2022/10/29/what-is-mlops-machine-learning-operations-why-do-you-need-mlops-for-machine-learning-and-deep-learning-projects/index.html)

- Arham Islam

[AI Hardware Accelerators For Machine Learning And Deep Learning \| How To Choose One](/content/2022/10/22/ai-hardware-accelerators-for-machine-learning-and-deep-learning-how-to-choose-one/index.html)

- Arham Islam

[Understanding the Role of Artificial Intelligence (AI) in Building Smart Cities and Top Startups Working on it](/content/2022/10/15/understanding-the-role-of-artificial-intelligence-ai-in-building-smart-cities-and-top-startups-working-on-it/index.html)

- Arham Islam

[Understanding The Artificial Intelligence (AI) Bill of Rights From The White House](/content/2022/10/06/understanding-the-artificial-intelligence-ai-bill-of-rights-from-the-white-house/index.html)

- Arham Islam

[Top Real World Applications of Reinforcement Learning in 2022](/content/2022/10/03/top-real-world-applications-of-reinforcement-learning-in-2022/index.html)

#### [RELATED ARTICLES](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/\#/index.html) [MORE FROM AUTHOR](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/\#/index.html)

### [How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)

### [Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order](/content/2026/06/13/anthropic-disables-claude-fable-5-and-mythos-5-after-us-government-order/ "Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order"/index.html)

### [Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)

### [A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)

### [Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard](/content/2026/06/12/google-releases-gemini-sql2-gemini-3-1-pro-text-to-sql-scores-80-04-on-bird-single-model-leaderboard/ "Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard"/index.html)

### [Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)

[prev-page](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/#/index.html)[next-page](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/#/index.html)

### [How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access,...](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 13, 2026[0](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/#respond/index.html)

In this tutorial, we implement a QwenPaw workflow that provides a practical environment for building and testing an agent-powered assistant. We install and initialize...

[Asif Razzaq](/content/author/6flvq/index.html)-June 13, 2026[0](/content/2026/06/13/anthropic-disables-claude-fable-5-and-mythos-5-after-us-government-order/#respond/index.html)

shutdown followed a US government export control directive citing national security authorities. All other Anthropic models, including Opus 4.8, remain available.

### [Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench...](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/#respond/index.html)

Moonshot AI has open-sourced Kimi K2.7-Code under a Modified MIT license. It is a coding-focused, agentic model built on Kimi K2.6, with a 256K context window and roughly 30% lower reasoning-token usage. Moonshot reports gains over K2.6 on six benchmarks, including +21.8% on Kimi Code Bench v2. The model is available via the Kimi API and Kimi Code.

### [A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph,...](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 12, 2026[0](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/#respond/index.html)

We build an end-to-end spatial graph learning pipeline using city2graph. We collect urban POI and street network data from OpenStreetMap, with a synthetic fallback for reliability. We engineer spatial features, construct several proximity graph families, and compare how each represents the same urban environment. We then build heterogeneous and homogeneous graphs, convert them to PyTorch Geometric, and train a GraphSAGE model to predict POI categories from spatial structure.

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/google-releases-gemini-sql2-gemini-3-1-pro-text-to-sql-scores-80-04-on-bird-single-model-leaderboard/#respond/index.html)

We look at Gemini-SQL2, the text-to-SQL capability Google Research announced on June 12, 2026. Powered by Gemini 3.1 Pro, it posted 80.04% execution accuracy on the BIRD single-model leaderboard. We explain what the score measures, how the leaderboard stacks up, and what Google has not yet disclosed. We also cover use cases and a schema-grounded implementation pattern.

### [Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6...](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/#respond/index.html)

Moonshot AI's Kimi Work is a local desktop agent for macOS and Windows. It runs a 300-sub-agent swarm, drives your logged-in browser via WebBridge, and schedules background jobs.

### [Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order...](/content/2026/06/12/zyphra-release-zamba2-vl-hybrid-mamba2-transformer-vision-language-models-that-cut-time-to-first-token-by-about-an-order-of-magnitude/ "Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/zyphra-release-zamba2-vl-hybrid-mamba2-transformer-vision-language-models-that-cut-time-to-first-token-by-about-an-order-of-magnitude/#respond/index.html)

Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about an order of magnitude.

### [A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical...](/content/2026/06/12/a-coding-implementation-on-monai-for-end-to-end-3d-spleen-segmentation-using-unet-on-medical-ct-volumes/ "A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical CT Volumes"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 12, 2026[0](/content/2026/06/12/a-coding-implementation-on-monai-for-end-to-end-3d-spleen-segmentation-using-unet-on-medical-ct-volumes/#respond/index.html)

In this tutorial, we build an end-to-end 3D medical image segmentation pipeline using MONAI to segment the spleen on the Medical Segmentation Decathlon Task09...

### [Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For...](/content/2026/06/11/perplexity-moves-deep-research-into-computer-routing-research-subtasks-across-20-frontier-models-for-reports-decks-and-dashboards/ "Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For Reports, Decks, And Dashboards"/index.html)

[Michal Sutter](/content/author/michal-sutter/index.html)-June 11, 2026[0](/content/2026/06/11/perplexity-moves-deep-research-into-computer-routing-research-subtasks-across-20-frontier-models-for-reports-decks-and-dashboards/#respond/index.html)

Deep Research now lives inside Perplexity Computer, breaking hard questions into subtasks and routing across 20+ frontier models.

### [xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and...](/content/2026/06/11/xai-ships-grok-build-plugin-marketplace-with-mongodb-vercel-sentry-chrome-devtools-cloudflare-and-superpowers-plugins-at-launch/ "xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and Superpowers Plugins at Launch"/index.html)

[Michal Sutter](/content/author/michal-sutter/index.html)-June 11, 2026[0](/content/2026/06/11/xai-ships-grok-build-plugin-marketplace-with-mongodb-vercel-sentry-chrome-devtools-cloudflare-and-superpowers-plugins-at-launch/#respond/index.html)

Grok Build's in-terminal marketplace bundles skills, agents, hooks, and MCP servers, with commit-SHA verification on every remote plugin.

- [miniCON Event 2025](https://pxl.to/hki7r39)
- [Download](/content/download/index.html)
  - [AI Magazine/Report](/content/ai-magazine/index.html)
- [Privacy & TC](/content/privacy-policy/index.html)
- [Cookie Policy](/content/cookie-policy/index.html)
- [Newsletter](https://www.aidevsignals.com/)
- [Partnership and Promotion](https://forms.gle/mjneG2kKPjDu6Hv8A)

© Copyright Reserved @2025 Marktechpost AI Media Inc

[Toggle photo metadata visibility](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/#/index.html)[Toggle photo comments visibility](/content/2026/03/21/safely-deploying-ml-models-to-production-four-controlled-strategies-a-b-canary-interleaved-shadow-testing/#/index.html)

Loading Comments...

Write a Comment...

Email (Required)Name (Required)Website
