We use cookies on our website to give you the most relevant experience by remembering your preferences and repeat visits. By clicking “Accept”, you consent to the use of ALL the cookies. .

Cookie settingsACCEPT

NecessaryAlways Active

Necessary cookies are absolutely essential for the website to function properly. This category only includes cookies that ensures basic functionalities and security features of the website. These cookies do not store any personal information.

- Cookie

\_\_cf\_bm

- Duration

1 hour

- Description

This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.

- Cookie

\_pxvid

- Duration

1 year

- Description

PerimeterX sets this cookie to detect fraud and bot activity.

- Cookie

\_px3

- Duration

6 minutes

- Description

This cookie is set by the Bloomberg to protect the site from BOT attacks.

- Cookie

CookieLawInfoConsent

- Duration

1 year

- Description

CookieYes sets this cookie to record the default button state of the corresponding category and the status of CCPA. It works only in coordination with the primary cookie.

- Cookie

cookielawinfo-checkbox-necessary

- Duration

11 months

- Description

This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".

- Cookie

cookielawinfo-checkbox-others

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie stores user consent for cookies in the category "Others".

- Cookie

cookielawinfo-checkbox-non-necessary

- Duration

11 months

- Description

This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Non Necessary".

- Cookie

cookielawinfo-checkbox-analytics

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Analytics" category.

- Cookie

cookielawinfo-checkbox-performance

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie stores the user consent for cookies in the category "Performance".

- Cookie

cookielawinfo-checkbox-uncategorized

- Duration

1 year

- Description

The cookie is set by the GDPR Cookie Consent plugin to record the user consent for cookies in the category "Uncategorized".

- Cookie

cookielawinfo-checkbox-functional

- Duration

1 year

- Description

The GDPR Cookie Consent plugin sets the cookie to record the user consent for the cookies in the category "Functional".

- Cookie

cookielawinfo-checkbox-advertisement

- Duration

1 year

- Description

Set by the GDPR Cookie Consent plugin, this cookie records the user consent for the cookies in the "Advertisement" category.

- Cookie

wpEmojiSettingsSupports

- Duration

session

- Description

WordPress sets this cookie when a user interacts with emojis on a WordPress site. It helps determine if the user's browser can display emojis properly.

- Cookie

VISITOR\_PRIVACY\_METADATA

- Duration

6 months

- Description

YouTube sets this cookie to store the user's cookie consent state for the current domain.

- Cookie

viewed\_cookie\_policy

- Duration

11 months

- Description

The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.

- Cookie

PHPSESSID

- Duration

- Description

This cookie is native to PHP applications. The cookie is used to store and identify a users' unique session ID for the purpose of managing user session on the website. The cookie is a session cookies and is deleted when all the browser windows are closed.

- Cookie

\_\_cfduid

- Duration

4 weeks

- Description

The cookie is set by CloudFare. The cookie is used to identify individual clients behind a shared IP address d apply security settings on a per-client basis. It doesnot correspond to any user ID in the web application and does not store any personally identifiable information.

Functional

Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.

- Cookie

yt-remote-connected-devices

- Duration

never

- Description

YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.

- Cookie

ytidb::LAST\_RESULT\_ENTRY\_KEY

- Duration

never

- Description

The cookie ytidb::LAST\_RESULT\_ENTRY\_KEY is used by YouTube to store the last search result entry that was clicked by the user. This information is used to improve the user experience by providing more relevant search results in the future.

- Cookie

yt-remote-device-id

- Duration

never

- Description

YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.

- Cookie

yt-remote-session-name

- Duration

session

- Description

The yt-remote-session-name cookie is used by YouTube to store the user's video player preferences using embedded YouTube video.

- Cookie

yt-remote-fast-check-period

- Duration

session

- Description

The yt-remote-fast-check-period cookie is used by YouTube to store the user's video player preferences for embedded YouTube videos.

- Cookie

yt-remote-session-app

- Duration

session

- Description

The yt-remote-session-app cookie is used by YouTube to store user preferences and information about the interface of the embedded YouTube video player.

- Cookie

yt-remote-cast-available

- Duration

session

- Description

The yt-remote-cast-available cookie is used to store the user's preferences regarding whether casting is available on their YouTube video player.

- Cookie

yt-remote-cast-installed

- Duration

session

- Description

The yt-remote-cast-installed cookie is used to store the user's video player preferences using embedded YouTube video.

- Cookie

na\_id

- Duration

1 year

- Description

This cookie is set by Addthis.com to enable sharing of links on social media platforms like Facebook and Twitter

- Cookie

vc

- Duration

1 year

- Description

This cookie is set by addthis.com on sites that allow sharing on social media.

- Cookie

\_\_atuvc

- Duration

1 year

- Description

This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.

- Cookie

\_\_atuvs

- Duration

30 minutes

- Description

This cookie is set by Addthis to make sure you see the updated count if you share a page and return to it before our share count cache is updated.

- Cookie

ouid

- Duration

1 year

- Description

The cookie is set by Addthis which enables the content of the website to be shared across different networking and social sharing websites.

Analytics

Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.

- Cookie

\_ga\_\*

- Duration

1 year 1 month 4 days

- Description

Google Analytics sets this cookie to store and count page views.

- Cookie

\_ga

- Duration

2 years

- Description

This cookie is installed by Google Analytics. The cookie is used to calculate visitor, session, camapign data and keep track of site usage for the site's analytics report. The cookies store information anonymously and assigns a randoly generated number to identify unique visitors.

- Cookie

sbjs\_migrations

- Duration

session

- Description

Sourcebuster sets this cookie to identify the source of a visit and stores user action information in cookies. This analytical and behavioural cookie is used to enhance the visitor experience on the website.

- Cookie

sbjs\_current\_add

- Duration

session

- Description

- Cookie

sbjs\_first\_add

- Duration

session

- Description

- Cookie

sbjs\_current

- Duration

session

- Description

- Cookie

sbjs\_first

- Duration

session

- Description

- Cookie

sbjs\_udata

- Duration

session

- Description

- Cookie

sbjs\_session

- Duration

1 hour

- Description

- Cookie

tk\_or

- Duration

1 year 1 month 4 days

- Description

JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.

- Cookie

tk\_r3d

- Duration

3 days

- Description

JetPack installs this cookie to collect internal metrics for user activity and improve user experience.

- Cookie

tk\_lr

- Duration

1 year

- Description

JetPack plugin sets this referral cookie on sites using WooCommerce, which analyzes referrer behaviour for Jetpack.

- Cookie

tk\_ai

- Duration

1 year

- Description

JetPack sets this cookie to store a randomly-generated anonymous ID used only within the admin area and for general analytics tracking.

- Cookie

tk\_tc

- Duration

session

- Description

JetPack sets this cookie to record details on how users use the website.

- Cookie

\_gat\_gtag\_UA\_5784146\_31

- Duration

1 minute

- Description

Google Used to distinguish users.

- Cookie

GPS

- Duration

30 minutes

- Description

This cookie is set by Youtube and registers a unique ID for tracking users based on their geographical location

- Cookie

\_\_gads

- Duration

2 years

- Description

This cookie is set by Google and stored under the name dounleclick.com. This cookie is used to track how many times users see a particular advert which helps in measuring the success of the campaign and calculate the revenue generated by the campaign. These cookies can only be read from the domain that it is set on so it will not track any data while browsing through another sites.

- Cookie

uvc

- Duration

1 year

- Description

The cookie is set by addthis.com to determine the usage of Addthis.com service.

- Cookie

ad-id

- Duration

7 months

- Description

Provided by amazon-adsystem.com for tracking user actions on other websites to provide targeted content

- Cookie

\_gat\_gtag\_UA\_116563943\_1

- Duration

1 minute

- Description

Google uses this cookie to distinguish users.

- Cookie

\_gid

- Duration

1 day

- Description

This cookie is installed by Google Analytics. The cookie is used to store information of how visitors use a website and helps in creating an analytics report of how the wbsite is doing. The data collected including the number visitors, the source where they have come from, and the pages viisted in an anonymous form.

Performance

Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.

- Cookie

YSC

- Duration

- Description

This cookies is set by Youtube and is used to track the views of embedded videos.

- Cookie

\_gat

- Duration

1 minute

- Description

This cookies is installed by Google Universal Analytics to throttle the request rate to limit the colllection of data on high traffic sites.

Advertisement

Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.

- Cookie

COMPASS

- Duration

1 hour

- Description

The COMPASS cookie is used by Yahoo to deliver targeted advertising based on user's online behavior.

- Cookie

NID

- Duration

5 months

- Description

This cookie is used to a profile based on user's interest and display personalized ads to the users.

- Cookie

\_\_Secure-YNID

- Duration

6 months

- Description

Google cookie used to protect user security and prevent fraud, especially during the login process.

- Cookie

\_\_Secure-ROLLOUT\_TOKEN

- Duration

6 months

- Description

YouTube sets this cookie to manage feature rollout and experimentation. It helps Google control which new features or interface changes are shown to users as part of testing and staged rollouts, ensuring consistent experience for a given user during an experiment.

- Cookie

yt.innertube::nextId

- Duration

never

- Description

YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.

- Cookie

yt.innertube::requests

- Duration

never

- Description

YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.

- Cookie

VISITOR\_INFO1\_LIVE

- Duration

5 months

- Description

This cookie is set by Youtube. Used to track the information of the embedded YouTube videos on a website.

- Cookie

TapAd\_TS

- Duration

1 month

- Description

The cookie is set by Tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising.

- Cookie

TapAd\_DID

- Duration

1 month

- Description

The cookie is set by tapad.com. The purpose of the cookie is to track users across devices to enable targeted advertising

- Cookie

personalization\_id

- Duration

2 years

- Description

This cookie is set by twitter.com. It is used integrate the sharing features of this social media. It also stores information about how the user uses the website for tracking and targeting.

- Cookie

uid

- Duration

1 year

- Description

This cookie is used to measure the number and behavior of the visitors to the website anonymously. The data includes the number of visits, average duration of the visit on the website, pages visited, etc. for the purpose of better understanding user preferences for targeted advertisments.

- Cookie

loc

- Duration

1 year

- Description

This cookie is set by Addthis. This is a geolocation cookie to understand where the users sharing the information are located.

- Cookie

IDE

- Duration

2 years

- Description

Used by Google DoubleClick and stores information about how the user uses the website and any other advertisement before visiting the website. This is used to present users with ads that are relevant to them according to the user profile.

- Cookie

di2

- Duration

1 year

- Description

This cookie is set by addthis.com on sites that allows sharing on social media. The cookie is used to track user behavior anonymously to generate usage trends to improve relevance to their services and advertising.

Others

Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.

- Cookie

pxcts

- Duration

session

- Description

Description is currently not available.

- Cookie

\_pxttld

- Duration

session

- Description

Description is currently not available.

- Cookie

SGPBShowingLimitationDomain77659

- Duration

2 days

- Description

Description is currently not available.

- Cookie

\_\_Secure-YEC

- Duration

past

- Description

YouTube sets this cookie to stores the user's video player preferences using embedded YouTube video

- Cookie

S

- Duration

1 hour

- Description

Used by Yahoo to provide ads, content or analytics.

- Cookie

test\_cookie

- Duration

11 months

- Description

This cookie is set by doubleclick.net. The purpose of the cookie is to determine if the users' browser supports cookies.

- Cookie

sc\_at

- Duration

1 year

- Description

Snapchat sets this cookie for showing relevant advertising based on the user’s movement.

- Cookie

TapAd\_3WAY\_SYNCS

- Duration

1 month

- Description

TapAd sets this cookie for data synchronization with advertising networks.

- Cookie

\_pin\_unauth

- Duration

1 year

- Description

Pinterest set this cookie to group actions for users who cannot be identified.

- Cookie

sc\_anonymous\_id

- Duration

9 years

- Description

Soundcloud sets this cookie to enable visitors to embed content or files on the website.

- Cookie

um

- Duration

1 year

- Description

Set by addthis.com.(Purpose not known)

- Cookie

DCRP\_Categories

- Duration

4 weeks

- Description

Description is currently not available.

- Cookie

vuid

- Duration

2 years

- Description

Vimeo installs this cookie to collect tracking information by setting a unique ID to embed videos on the website.

- Cookie

X-AB

- Duration

1 day

- Description

Adobe Analytics sets this cookie in context with multi-variate testing. This is a tool used to combine or change content on the website. This allows the website to find the best variation or edition of the site.

- Cookie

YTC

- Duration

10 minutes

- Description

YouTube sets the YTC cookie to manage the embed and viewing of videos on the website.

- Cookie

sp\_t

- Duration

1 month

- Description

The sp\_t cookie is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.

- Cookie

sp\_landing

- Duration

1 day

- Description

The sp\_landing is set by Spotify to implement audio content from Spotify on the website and also registers information on user interaction related to the audio content.

- Cookie

\_\_asc

- Duration

30 minutes

- Description

Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.

- Cookie

\_\_auc

- Duration

1 year

- Description

Alexa Metrics sets this cookie to track and report information to the Alexa analytics service.

- Cookie

AWSESS

- Duration

- Description

Awin sets this to ensure the same kind of advertisement is not shown to the user.

- Cookie

nevercache-b39818

- Duration

session

- Description

Description is currently not available.

REJECTSave My PreferencesACCEPT

Powered by

NewsHub](/content/site-root.html)

[Premium Content](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/# "Premium Content"/index.html)

[Read our exclusive articles](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/# "Read our exclusive articles"/index.html)

[Facebook](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/# "Facebook"/index.html)

[Instagram](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/# "Instagram"/index.html)

[X](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/# "X"/index.html)

[Discord](https://pxl.to/ivxz41s "Discord")[Linkedin](https://www.linkedin.com/company/marktechpost/?viewAsMember=true "Linkedin")[Reddit](https://www.reddit.com/r/machinelearningnews/ "Reddit")[X](https://twitter.com/Marktechpost "X")

- [Home](/content/site-root.html)
- [Open Source/Weights](/content/category/technology/open-source/index.html)
- [AI Agents](/content/category/editors-pick/ai-agents/index.html)
- [Tutorials](/content/category/tutorials/index.html)
- [Voice AI](/content/category/technology/artificial-intelligence/voice-ai/index.html)
- [Robotics](/content/category/robotics/index.html)
- [Newsletter](https://www.aidevsignals.com/)
- [→ Partner with Us](https://forms.gle/CY1eqZzuWFQBp7dH9)

Search

NewsHub](/content/site-root.html)

NewsHub](/content/site-root.html)

Search

[Home](/content/ ""/index.html)[Tech News](/content/category/tech-news/ "View all posts in Tech News"/index.html)[AI Paper Summary](/content/category/tech-news/ai-paper-summary/ "View all posts in AI Paper Summary"/index.html)LEAN-GitHub: A Large-Scale Dataset for Advancing Automated Theorem Proving

[tinyfish.aiOpen Source\\
\\
Big **Set**\\
\\
Describe your ideal dataset in plain English, and BigSet builds it.\\
\\
dataset.build()auto·refresh\\
\\
✓\\
\\
✓\\
\\
✓\\
\\
✓\\
\\
Explore on GitHub→](https://pxllnk.co/cuv4rk8)

- [Tech News](/content/category/tech-news/index.html)
- [AI Paper Summary](/content/category/tech-news/ai-paper-summary/index.html)
- [Technology](/content/category/technology/index.html)
- [AI Shorts](/content/category/technology/ai-shorts/index.html)
- [Artificial Intelligence](/content/category/technology/artificial-intelligence/index.html)
- [Editors Pick](/content/category/editors-pick/index.html)
- [Staff](/content/category/editors-pick/staff/index.html)

[Add as a preferred\\
\\
source on Google](https://www.google.com/preferences/source?q=https://www.marktechpost.com/)

Theorem proving in mathematics faces growing challenges due to increasing proof complexity. Formalized systems like Lean, Isabelle, and Coq offer computer-verifiable proofs, but creating these demands substantial human effort. Large language models (LLMs) show promise in solving high-school-level math problems using proof assistants, yet their performance still needs to improve due to data scarcity. Formal languages require significant expertise, resulting in limited corpora. Unlike conventional programming languages, formal proof languages contain hidden intermediate information, making raw language corpora unsuitable for training. This scarcity persists despite the existence of valuable human-written corpora. Auto-formalization efforts, while helpful, cannot fully substitute human-crafted data in quality and diversity.

Existing attempts to address theorem-proving challenges have evolved significantly with modern proof assistants like Coq, Isabelle, and Lean having expanded formal systems beyond first-order logic, increasing interest in automated theorem proving (ATP). The recent integration of large language models has further advanced this field. Early ATP approaches used traditional methods like KNN or GNN, with some employing reinforcement learning. Recent efforts utilize deep transformer-based methods, treating theorems as plain text. Many learning-based systems (e.g., GPT-f, PACT, Llemma) train language models on (proof state, next-tactic) pairs and use tree search for theorem proving. Alternative approaches involve LLMs generating entire proofs independently or based on human-provided proofs. Data extraction tools are crucial for ATP, capturing intermediate states invisible in code but visible during runtime. Tools exist for various proof assistants, but Lean 4 tools face challenges in massive extraction across multiple projects due to single-project design limitations. Some methods also explore incorporating informal proofs into formal proofs, broadening the scope of ATP research.

Researchers from The Chinese University of Hong Kong propose **_LEAN-GitHub_**, a large-scale Lean dataset that complements the well-utilized Mathlib dataset. This innovative approach provides an open-source Lean repositories on GitHub, significantly expanding the available data for training theorem-proving models. The researchers developed a scalable pipeline to enhance extraction efficiency and parallelism, enabling the exploitation of valuable data from previously uncompiled and unextracted Lean corpus. Also, they provide a solution to the state duplication problem common in tree-proof search methods.

The LEAN-GitHub dataset construction process involved several key steps and innovations:

1. Repository Selection: The researchers identified 237 Lean 4 repositories  (GitHub does not differentiate between Lean 3 and Lean 4) on GitHub, estimating approximately 48,091 theorems. After discarding 90 repositories with deprecated Lean 4 versions, 147 remained. Only 61 of these could be compiled without modifications.
2. Compilation Challenges: The team developed automated scripts to find the closest official releases for projects using non-official Lean 4 versions. They also addressed the issue of isolated files within empty Lean projects.
3. Source Code Compilation: Instead of using the Lake tool, they called the Leanc compiler directly. This approach allowed for compiling non-compliant Lean projects and isolated files, which Lake couldn’t handle. They extended Lake’s import graph and created a custom compiling script with increased parallelism.
4. Extraction Process: Building upon LeanDojo, the team implemented data extraction for isolated files and restructured the implementation to increase parallelism. This approach overcame bottlenecks in network connection and computational redundancies.
5. Results: Out of 8,639 Lean source files, 6,352 and 42,000 theorems were successfully extracted. The final dataset includes 2,133 files and 28,000 theorems with valid tactic information.

The resulting LEAN-GitHub dataset is diverse, covering various mathematical fields including logic, first-order logic, matroid theory, and arithmetic. It contains cutting-edge mathematical topics, data structures, and Olympiad-level problems. Compared to existing datasets, LEAN-GitHub offers a unique combination of human-written content, intermediate states, and diverse complexity levels, making it a valuable resource for advancing automated theorem proving and formal mathematics.

InternLM2-StepProver, trained on the diverse LEAN-GitHub dataset, demonstrates exceptional formal reasoning abilities across various benchmarks. It achieves state-of-the-art performance on miniF2F (63.9% on Valid, 54.5% on Test), surpassing previous models. On ProofNet, it attains an 18.1% Pass@1 rate, outperforming the previous leader. For [PutnamBench](https://github.com/trishullab/PutnamBench), it solves 5 problems in a single pass, including the previously unsolved [Putnam](https://en.wikipedia.org/wiki/William_Lowell_Putnam_Mathematical_Competition) 1988 B2. These results span high-school to advanced undergraduate-level mathematics, showcasing InternLM2-StepProver’s versatility and the effectiveness of the LEAN-GitHub dataset in training advanced theorem-proving models.

LEAN-GitHub, a large-scale dataset extracted from open Lean 4 repositories, contains 28,597 theorems and 218,866 tactics. This diverse dataset was used to train InternLM2-StepProver, achieving state-of-the-art performance in Lean 4 formal reasoning. Models trained on LEAN-GitHub demonstrate improved performance across various mathematical fields and difficulty levels, highlighting the dataset’s effectiveness in enhancing reasoning capabilities. By open-sourcing LEAN-GitHub, the researchers aim to help the community better utilize under-exploited information in raw corpora and advance mathematical reasoning. This contribution could significantly accelerate progress in automated theorem proving and formal mathematics.

* * *

Check out the **[Paper](https://arxiv.org/abs/2407.17227)** and [**Dataset**.](https://huggingface.co/datasets/internlm/Lean-Github) All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on **[Twitter](https://twitter.com/Marktechpost)** and join our **[Telegram Channel](https://pxl.to/at72b5j)** and [**LinkedIn Gr**](https://www.linkedin.com/groups/13668564/) [**oup**](https://www.linkedin.com/groups/13668564/). **If you like our work, you will love our** [**newsletter..**](https://marktechpost-newsletter.beehiiv.com/subscribe)

Don’t Forget to join our **[47k+ ML SubReddit](https://www.reddit.com/r/machinelearningnews/)**

**Find Upcoming [AI Webinars here](/content/ai-webinars-list-llms-rag-generative-ai-ml-vector-database/index.html)**

##### [Mohammad Asjad](/content/author/mohammad_asjad/index.html)

[\+ postsBio](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/#/index.html)

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

- Mohammad Asjad

[Agentic AI in Financial Services: IBM’s Whitepaper Maps Opportunities, Risks, and Responsible Integration](/content/2025/05/19/agentic-ai-in-financial-services-ibms-whitepaper-maps-opportunities-risks-and-responsible-integration/index.html)

- Mohammad Asjad

[Critical Security Vulnerabilities in the Model Context Protocol (MCP): How Malicious Tools and Deceptive Contexts Exploit AI Agents](/content/2025/05/18/critical-security-vulnerabilities-in-the-model-context-protocol-mcp-how-malicious-tools-and-deceptive-contexts-exploit-ai-agents/index.html)

- Mohammad Asjad

[Stability AI Introduces Adversarial Relativistic-Contrastive (ARC) Post-Training and Stable Audio Open Small: A Distillation-Free Breakthrough for Fast, Diverse, and Efficient Text-to-Audio Generation Across Devices](/content/2025/05/15/stability-ai-introduces-adversarial-relativistic-contrastive-arc-post-training-and-stable-audio-open-small-a-distillation-free-breakthrough-for-fast-diverse-and-efficient-text-to-audio-generation/index.html)

- Mohammad Asjad

[Meta AI Introduces CATransformers: A Carbon-Aware Machine Learning Framework to Co-Optimize AI Models and Hardware for Sustainable Edge Deployment](/content/2025/05/14/meta-ai-introduces-catransformers-a-carbon-aware-machine-learning-framework-to-co-optimize-ai-models-and-hardware-for-sustainable-edge-deployment/index.html)

- Mohammad Asjad

[Enterprise AI Without GPU Burn: Salesforce’s xGen-small Optimizes for Context, Cost, and Privacy](/content/2025/05/09/enterprise-ai-without-gpu-burn-salesforces-xgen-small-optimizes-for-context-cost-and-privacy/index.html)

- Mohammad Asjad

[ServiceNow AI Released Apriel-Nemotron-15b-Thinker: A Compact Yet Powerful Reasoning Model Optimized for Enterprise-Scale Deployment and Efficiency](/content/2025/05/09/servicenow-ai-released-apriel-nemotron-15b-thinker-a-compact-yet-powerful-reasoning-model-optimized-for-enterprise-scale-deployment-and-efficiency/index.html)

- Mohammad Asjad

[Researchers from Fudan University Introduce Lorsa: A Sparse Attention Mechanism That Recovers Atomic Attention Units Hidden in Transformer Superposition](/content/2025/05/07/researchers-from-fudan-university-introduce-lorsa-a-sparse-attention-mechanism-that-recovers-atomic-attention-units-hidden-in-transformer-superposition/index.html)

- Mohammad Asjad

[A Step-by-Step Guide to Implement Intelligent Request Routing with Claude](/content/2025/05/06/a-step-by-step-guide-to-implement-intelligent-request-routing-with-claude/index.html)

- Mohammad Asjad

[Scaling Reinforcement Learning Beyond Math: Researchers from NVIDIA AI and CMU Propose Nemotron-CrossThink for Multi-Domain Reasoning with Verifiable Reward Modeling](/content/2025/05/04/scaling-reinforcement-learning-beyond-math-researchers-from-nvidia-ai-and-cmu-propose-nemotron-crossthink-for-multi-domain-reasoning-with-verifiable-reward-modeling/index.html)

- Mohammad Asjad

[Vision Foundation Models: Implementation and Business Applications](/content/2025/05/03/vision-foundation-models-implementation-and-business-applications/index.html)

- Mohammad Asjad

[LLMs Can Now Reason in Parallel: UC Berkeley and UCSF Researchers Introduce Adaptive Parallel Reasoning to Scale Inference Efficiently Without Exceeding Context Windows](/content/2025/05/02/llms-can-now-reason-in-parallel-uc-berkeley-and-ucsf-researchers-introduce-adaptive-parallel-reasoning-to-scale-inference-efficiently-without-exceeding-context-windows/index.html)

- Mohammad Asjad

[Training LLM Agents Just Got More Stable: Researchers Introduce StarPO-S and RAGEN to Tackle Multi-Turn Reasoning and Collapse in Reinforcement Learning](/content/2025/05/01/training-llm-agents-just-got-more-stable-researchers-introduce-starpo-s-and-ragen-to-tackle-multi-turn-reasoning-and-collapse-in-reinforcement-learning/index.html)

- Mohammad Asjad

[The WAVLab Team Releases of VERSA: A Comprehensive and Versatile Evaluation Toolkit for Assessing Speech, Audio, and Music Signals](/content/2025/04/28/the-wavlab-team-is-releases-of-versa-a-comprehensive-and-versatile-evaluation-toolkit-for-assessing-speech-audio-and-music-signals/index.html)

- Mohammad Asjad

[Google DeepMind Research Introduces QuestBench: Evaluating LLMs’ Ability to Identify Missing Information in Reasoning Tasks](/content/2025/04/25/google-deepmind-research-introduces-questbench-evaluating-llms-ability-to-identify-missing-information-in-reasoning-tasks/index.html)

- Mohammad Asjad

[LLMs Can Now Solve Challenging Math Problems with Minimal Data: Researchers from UC Berkeley and Ai2 Unveil a Fine-Tuning Recipe That Unlocks Mathematical Reasoning Across Difficulty Levels](/content/2025/04/18/llms-can-now-solve-challenging-math-problems-with-minimal-data-researchers-from-uc-berkeley-and-ai2-unveil-a-fine-tuning-recipe-that-unlocks-mathematical-reasoning-across-difficulty-levels/index.html)

- Mohammad Asjad

[LLM Reasoning Benchmarks are Statistically Fragile: New Study Shows Reinforcement Learning RL Gains often Fall within Random Variance](/content/2025/04/15/llm-reasoning-benchmarks-are-statistically-fragile-new-study-shows-reinforcement-learning-rl-gains-often-fall-within-random-variance/index.html)

- Mohammad Asjad

[Multimodal Models Don’t Need Late Fusion: Apple Researchers Show Early-Fusion Architectures are more Scalable, Efficient, and Modality-Agnostic](/content/2025/04/14/multimodal-models-dont-need-late-fusion-apple-researchers-show-early-fusion-architectures-are-more-scalable-efficient-and-modality-agnostic/index.html)

- Mohammad Asjad

[Step by Step Coding Guide to Build a Neural Collaborative Filtering (NCF) Recommendation System with PyTorch](/content/2025/04/11/step-by-step-coding-guide-to-build-a-neural-collaborative-filtering-ncf-recommendation-system-with-pytorch/index.html)

- Mohammad Asjad

[This AI Paper Introduces a Machine Learning Framework to Estimate the Inference Budget for Self-Consistency and GenRMs (Generative Reward Models)](/content/2025/04/10/this-ai-paper-introduces-a-machine-learning-framework-to-estimate-the-inference-budget-for-self-consistency-and-genrms-generative-reward-models/index.html)

- Mohammad Asjad

[MMSearch-R1: End-to-End Reinforcement Learning for Active Image Search in LMMs](/content/2025/04/06/mmsearch-r1-end-to-end-reinforcement-learning-for-active-image-search-in-lmms/index.html)

- Mohammad Asjad

[Anthropic’s Evaluation of Chain-of-Thought Faithfulness: Investigating Hidden Reasoning, Reward Hacks, and the Limitations of Verbal AI Transparency in Reasoning Models](/content/2025/04/05/anthropics-evaluation-of-chain-of-thought-faithfulness-investigating-hidden-reasoning-reward-hacks-and-the-limitations-of-verbal-ai-transparency-in-reasoning-models/index.html)

- Mohammad Asjad

[Building Your AI Q&A Bot for Webpages Using Open Source AI Models](/content/2025/04/04/building-your-ai-qa-bot-for-webpages-using-open-source-ai-models/index.html)

- Mohammad Asjad

[DeltaProduct: An AI Method that Balances Expressivity and Efficiency of the Recurrence Computation, Improving State-Tracking in Linear Recurrent Neural Networks](/content/2025/04/01/deltaproduct-an-ai-method-that-balances-expressivity-and-efficiency-of-the-recurrence-computation-improving-state-tracking-in-linear-recurrent-neural-networks/index.html)

- Mohammad Asjad

[PydanticAI: Advancing Generative AI Agent Development through Intelligent Framework Design](/content/2025/03/25/pydanticai-advancing-generative-ai-agent-development-through-intelligent-framework-design/index.html)

- Mohammad Asjad

[TxAgent: An AI Agent that Delivers Evidence-Grounded Treatment Recommendations by Combining Multi-Step Reasoning with Real-Time Biomedical Tool Integration](/content/2025/03/23/txagent-an-ai-agent-that-delivers-evidence-grounded-treatment-recommendations-by-combining-multi-step-reasoning-with-real-time-biomedical-tool-integration/index.html)

- Mohammad Asjad

[Building a Retrieval-Augmented Generation (RAG) System with FAISS and Open-Source LLMs](/content/2025/03/18/building-a-retrieval-augmented-generation-rag-system-with-faiss-and-open-source-llms/index.html)

- Mohammad Asjad

[Meet PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC](/content/2025/03/15/meet-pc-agent-a-hierarchical-multi-agent-collaboration-framework-for-complex-task-automation-on-pc/index.html)

- Mohammad Asjad

[Implementing Text-to-Speech TTS with BARK Using Hugging Face’s Transformers library in a Google Colab environment](/content/2025/03/11/implementing-text-to-speech-tts-with-bark-using-hugging-faces-transformers-library-in-a-google-colab-environment/index.html)

- Mohammad Asjad

[Salesforce AI Releases Text2Data: A Training Framework for Low-Resource Data Generation](/content/2025/03/09/salesforce-ai-releases-text2data-a-training-framework-for-low-resource-data-generation/index.html)

- Mohammad Asjad

[Q-Filters: A Training-Free AI Method for Efficient KV Cache Compression](/content/2025/03/06/q-filters-a-training-free-ai-method-for-efficient-kv-cache-compression/index.html)

- Mohammad Asjad

[Starter Guide For Running Large Language Models LLMs](/content/2025/03/06/starter-guide-for-running-large-language-models-llms/index.html)

- Mohammad Asjad

[Thinking Harder, Not Longer: Evaluating Reasoning Efficiency in Advanced Language Models](/content/2025/02/28/thinking-harder-not-longer-evaluating-reasoning-efficiency-in-advanced-language-models/index.html)

- Mohammad Asjad

[CoSyn: An AI Framework that Leverages the Coding Capabilities of Text-only Large Language Models (LLMs) to Automatically Create Synthetic Text-Rich Multimodal Data](/content/2025/02/25/cosyn-an-ai-framework-that-leverages-the-coding-capabilities-of-text-only-large-language-models-llms-to-automatically-create-synthetic-text-rich-multimodal-data/index.html)

- Mohammad Asjad

[Why Do Task Vectors Exist in Pretrained LLMs? This AI Research from MIT and Improbable AI Uncovers How Transformers Form Internal Abstractions and the Mechanisms Behind in-Context Learning (ICL)](/content/2024/12/23/why-do-task-vectors-exist-in-pretrained-llms-this-ai-research-from-mit-and-improbable-ai-uncovers-how-transformers-form-internal-abstractions-and-the-mechanisms-behind-in-context-learning-icl/index.html)

- Mohammad Asjad

[Scaling Language Model Evaluation: From Thousands to Millions of Tokens with BABILong](/content/2024/12/19/scaling-language-model-evaluation-from-thousands-to-millions-of-tokens-with-babilong/index.html)

- Mohammad Asjad

[The Role of Specifications in Modularizing Large Language Models](/content/2024/12/17/the-role-of-specifications-in-modularizing-large-language-models/index.html)

- Mohammad Asjad

[Google Released State of the Art ‘Veo 2’ for Video Generation and ‘Improved Imagen 3’ for Image Creation: Setting New Standards with 4K Video and Several Minutes Long Video Generation](/content/2024/12/17/google-released-state-of-the-art-veo-2-for-video-generation-and-improved-imagen-3-for-image-creation-setting-new-standards-with-4k-video-and-several-minutes-long-video-generation/index.html)

- Mohammad Asjad

[Meta FAIR Releases Meta Motivo: A New Behavioral Foundation Model for Controlling Virtual Physics-based Humanoid Agents for a Wide Range of Complex Whole-Body Tasks](/content/2024/12/16/meta-fair-releases-meta-motivo-a-new-behavioral-foundation-model-for-controlling-virtual-physics-based-humanoid-agents-for-a-wide-range-of-complex-whole-body-tasks/index.html)

- Mohammad Asjad

[Beyond the Mask: A Comprehensive Study of Discrete Diffusion Models](/content/2024/12/15/beyond-the-mask-a-comprehensive-study-of-discrete-diffusion-models/index.html)

- Mohammad Asjad

[Alibaba Qwen Researchers Introduced ProcessBench: A New AI Benchmark for Measuring the Ability to Identify Process Errors in Mathematical Reasoning](/content/2024/12/14/alibaba-qwen-researchers-introduced-processbench-a-new-ai-benchmark-for-measuring-the-ability-to-identify-process-errors-in-mathematical-reasoning/index.html)

- Mohammad Asjad

[Best-of-N Jailbreaking: A Multi-Modal AI Approach to Identifying Vulnerabilities in Large Language Models](/content/2024/12/13/best-of-n-jailbreaking-a-multi-modal-ai-approach-to-identifying-vulnerabilities-in-large-language-models/index.html)

- Mohammad Asjad

[Latent Functional Maps: A Robust Machine Learning Framework for Analyzing Neural Network Representations](/content/2024/12/10/latent-functional-maps-a-robust-machine-learning-framework-for-analyzing-neural-network-representations/index.html)

- Mohammad Asjad

[Voyage AI Introduces voyage-code-3: A New Next-Generation Embedding Model Optimized for Code Retrieval](/content/2024/12/09/voyage-ai-introduces-voyage-code-3-a-new-next-generation-embedding-model-optimized-for-code-retrieval/index.html)

- Mohammad Asjad

[Meet GRAPE: A Plug-and-Play Algorithm to Generalize Robot Policies via Preference Alignment](/content/2024/12/07/meet-grape-a-plug-and-play-algorithm-to-generalize-robot-policies-via-preference-alignment/index.html)

- Mohammad Asjad

[Top 20 Guardrails to Secure LLM Applications](/content/2024/12/06/top-20-guardrails-to-secure-llm-applications/index.html)

- Mohammad Asjad

[Allen Institute for AI: Open-Source Innovations with Ethical Commitments and Contributions in 2024](/content/2024/12/04/allen-institute-for-ai-open-source-innovations-with-ethical-commitments-and-contributions-in-2024/index.html)

- Mohammad Asjad

[Can You Turn Your Vision-Language Model from a Zero-Shot Model to Any-Shot Generalist? Meet LIxP, the Context-Aware Multimodal Framework](/content/2024/12/03/can-you-turn-your-vision-language-model-from-a-zero-shot-model-to-any-shot-generalist-meet-lixp-the-context-aware-multimodal-framework/index.html)

- Mohammad Asjad

[Characterizing and Mitigating Compute Express Link (CXL) Interference in Modern Memory Systems](/content/2024/12/03/characterizing-and-mitigating-compute-express-link-cxl-interference-in-modern-memory-systems/index.html)

- Mohammad Asjad

[ShowUI: A Vision-Language-Action Model for GUI Visual Agents that Addresses Key Challenges in UI Visual and Action Modeling](/content/2024/12/01/showui-a-vision-language-action-model-for-gui-visual-agents-that-addresses-key-challenges-in-ui-visual-and-action-modeling/index.html)

- Mohammad Asjad

[Geometry Distributions: Advancing Neural 3D Surface Modeling with Diffusion Models](/content/2024/11/30/geometry-distributions-advancing-neural-3d-surface-modeling-with-diffusion-models/index.html)

- Mohammad Asjad

[SEALONG: A Self-Improving AI Approach to Long-Context Reasoning in Large Language Models](/content/2024/11/28/sealong-a-self-improving-ai-approach-to-long-context-reasoning-in-large-language-models/index.html)

- Mohammad Asjad

[Quantum Neuromorphic Computing: Implementing Scalable Quantum Perceptrons](/content/2024/11/27/quantum-neuromorphic-computing-implementing-scalable-quantum-perceptrons/index.html)

- Mohammad Asjad

[Red Teaming for AI: Strengthening Safety and Trust through External Evaluation](/content/2024/11/25/red-teaming-for-ai-strengthening-safety-and-trust-through-external-evaluation/index.html)

- Mohammad Asjad

[KuaiFormer: A Transformer-Based Architecture for Large-Scale Short-Video Recommendation Systems](/content/2024/11/22/kuaiformer-a-transformer-based-architecture-for-large-scale-short-video-recommendation-systems/index.html)

- Mohammad Asjad

[Artificial Intelligence AI and Quantum Computing: Transforming Computational Frontiers](/content/2024/11/21/artificial-intelligence-ai-and-quantum-computing-transforming-computational-frontiers/index.html)

- Mohammad Asjad

[Deep Learning Meets Cybersecurity: A Hybrid Approach to Detecting DDoS Attacks with Unmatched Accuracy](/content/2024/11/20/deep-learning-meets-cybersecurity-a-hybrid-approach-to-detecting-ddos-attacks-with-unmatched-accuracy/index.html)

- Mohammad Asjad

[Stanford Researchers Propose ‘POSR’: A Unique AI Framework for Analyzing Educational Conversations Using Joint Segmentation and Retrieval](/content/2024/11/19/stanford-researchers-propose-posr-a-unique-ai-framework-for-analyzing-educational-conversations-using-joint-segmentation-and-retrieval/index.html)

- Mohammad Asjad

[H-DPO: Advancing Language Model Alignment through Entropy Control](/content/2024/11/17/h-dpo-advancing-language-model-alignment-through-entropy-control/index.html)

- Mohammad Asjad

[Meet OpenCoder: A Completely Open-Source Code LLM Built on the Transparent Data Process Pipeline and Reproducible Dataset](/content/2024/11/14/meet-opencoder-a-completely-open-source-code-llm-built-on-the-transparent-data-process-pipeline-and-reproducible-dataset/index.html)

- Mohammad Asjad

[Researchers from Snowflake and CMU Introduce SuffixDecoding: A Novel Model-Free Approach to Accelerating Large Language Model (LLM) Inference through Speculative Decoding](/content/2024/11/13/researchers-from-snowflake-and-cmu-introduce-suffixdecoding-a-novel-model-free-approach-to-accelerating-large-language-model-llm-inference-through-speculative-decoding/index.html)

- Mohammad Asjad

[The Semantic Hub: A Cognitive Approach to Language Model Representations](/content/2024/11/09/the-semantic-hub-a-cognitive-approach-to-language-model-representations/index.html)

- Mohammad Asjad

[Researchers from Stanford and Cornell Introduce APRICOT: A Novel AI Approach that Merges LLM-based Bayesian Active Preference Learning with Constraint-Aware Task Planning](/content/2024/11/07/researchers-from-stanford-and-cornell-introduce-apricot-a-novel-ai-approach-that-merges-llm-based-bayesian-active-preference-learning-with-constraint-aware-task-planning/index.html)

- Mohammad Asjad

[Nearest Neighbor Normalization: A Sublinear Approach to Improving Contrastive Retrieval](/content/2024/11/05/nearest-neighbor-normalization-a-sublinear-approach-to-improving-contrastive-retrieval/index.html)

- Mohammad Asjad

[Predicting and Interpreting In-Context Learning Curves Through Bayesian Scaling Laws](/content/2024/11/04/predicting-and-interpreting-in-context-learning-curves-through-bayesian-scaling-laws/index.html)

- Mohammad Asjad

[Multi-Scale Geometric Analysis of Language Model Features: From Atomic Patterns to Galaxy Structures](/content/2024/11/02/multi-scale-geometric-analysis-of-language-model-features-from-atomic-patterns-to-galaxy-structures/index.html)

- Mohammad Asjad

[AUTO-CEI: A Curriculum and Expert Iteration Approach to Elevate LLMs’ Response Precision and Control Refusal Rates Across Diverse Reasoning Domains](/content/2024/11/01/auto-cei-a-curriculum-and-expert-iteration-approach-to-elevate-llms-response-precision-and-control-refusal-rates-across-diverse-reasoning-domains/index.html)

- Mohammad Asjad

[CodeFavor: A Machine Learning Framework that Trains Pairwise Preference Models with Synthetic Code Preferences Generated from Code Evolution like Code Commits and Code Critiques](/content/2024/10/31/codefavor-a-machine-learning-framework-that-trains-pairwise-preference-models-with-synthetic-code-preferences-generated-from-code-evolution-like-code-commits-and-code-critiques/index.html)

- Mohammad Asjad

[SimpleToM: Evaluating Applied Theory of Mind Capabilities in Large Language Models](/content/2024/10/30/simpletom-evaluating-applied-theory-of-mind-capabilities-in-large-language-models/index.html)

- Mohammad Asjad

[LongRAG: A Robust RAG Framework for Long-Context Question Answering](/content/2024/10/28/longrag-a-robust-rag-framework-for-long-context-question-answering/index.html)

- Mohammad Asjad

[MiniCTX: Advancing Context-Dependent Theorem Proving in Large Language Models](/content/2024/10/27/minictx-advancing-context-dependent-theorem-proving-in-large-language-models/index.html)

- Mohammad Asjad

[Meta AI Researchers Introduce Token-Level Detective Reward Model (TLDR) to Provide Fine-Grained Annotations for Large Vision Language Models](/content/2024/10/26/meta-ai-researchers-introduce-token-level-detective-reward-model-tldr-to-provide-fine-grained-annotations-for-large-vision-language-models/index.html)

- Mohammad Asjad

[Multi-Scale Neural Audio Codec (SNAC): An Wxtension of Residual Vector Quantization that Uses Quantizers Operating at Multiple Temporal Resolutions](/content/2024/10/23/multi-scale-neural-audio-codec-snac-an-wxtension-of-residual-vector-quantization-that-uses-quantizers-operating-at-multiple-temporal-resolutions/index.html)

- Mohammad Asjad

[Scaling Diffusion transformers (DiT): An AI Framework for Optimizing Text-to-Image Models Across Compute Budgets](/content/2024/10/19/scaling-diffusion-transformers-dit-an-ai-framework-for-optimizing-text-to-image-models-across-compute-budgets/index.html)

- Mohammad Asjad

[Model Kinship: The Degree of Similarity or Relatedness between LLMs, Analogous to Biological Evolution](/content/2024/10/18/model-kinship-the-degree-of-similarity-or-relatedness-between-llms-analogous-to-biological-evolution/index.html)

- Mohammad Asjad

[AutoDAN-Turbo: A Black-Box Jailbreak Method for LLMs with a Lifelong Agent](/content/2024/10/16/autodan-turbo-a-black-box-jailbreak-method-for-llms-with-a-lifelong-agent/index.html)

- Mohammad Asjad

[Stochastic Prompt Construction for Effective In-Context Reinforcement Learning in Large Language Models](/content/2024/10/13/stochastic-prompt-construction-for-effective-in-context-reinforcement-learning-in-large-language-models/index.html)

- Mohammad Asjad

[UNC Chapel Hill Researchers Propose DataEnvGym: A Testbed of Teacher Environments for Data Generation Agents](/content/2024/10/12/unc-chapel-hill-researchers-propose-dataenvgym-a-testbed-of-teacher-environments-for-data-generation-agents/index.html)

- Mohammad Asjad

[ScienceAgentBench: A Rigorous AI Evaluation Framework for Language Agents in Scientific Discovery](/content/2024/10/11/scienceagentbench-a-rigorous-ai-evaluation-framework-for-language-agents-in-scientific-discovery/index.html)

- Mohammad Asjad

[SQ-LLaVA: A New Visual Instruction Tuning Method that Enhances General-Purpose Vision-Language Understanding and Image-Oriented Question Answering through Visual Self-Questioning](/content/2024/10/10/sq-llava-a-new-visual-instruction-tuning-method-that-enhances-general-purpose-vision-language-understanding-and-image-oriented-question-answering-through-visual-self-questioning/index.html)

- Mohammad Asjad

[From Prediction to Reasoning: Evaluating o1’s Impact on LLM Probabilistic Biases](/content/2024/10/08/from-prediction-to-reasoning-evaluating-o1s-impact-on-llm-probabilistic-biases/index.html)

- Mohammad Asjad

[NVIDIA AI Releases OpenMathInstruct-2: A Math Instruction Tuning Dataset with 14M Problem-Solution Pairs Generated Using the Llama3.1-405B-Instruct Model](/content/2024/10/07/nvidia-ai-releases-openmathinstruct-2-a-math-instruction-tuning-dataset-with-14m-problem-solution-pairs-generated-using-the-llama3-1-405b-instruct-model/index.html)

- Mohammad Asjad

[Exploring In-Context Reinforcement Learning in LLMs with Sparse Autoencoders](/content/2024/10/07/exploring-in-context-reinforcement-learning-in-llms-with-sparse-autoencoders/index.html)

- Mohammad Asjad

[AI-Assisted Causal Inference: Using LLMs to Revolutionize Instrumental Variable Selection](/content/2024/10/06/ai-assisted-causal-inference-using-llms-to-revolutionize-instrumental-variable-selection/index.html)

- Mohammad Asjad

[GemFilter: A Novel AI Approach to Accelerate LLM Inference and Reduce Memory Consumption for Long Context Inputs](/content/2024/10/05/gemfilter-a-novel-ai-approach-to-accelerate-llm-inference-and-reduce-memory-consumption-for-long-context-inputs/index.html)

- Mohammad Asjad

[The Impact of AI Chatbots on False Memory Formation: A Comprehensive Study](/content/2024/10/04/the-impact-of-ai-chatbots-on-false-memory-formation-a-comprehensive-study/index.html)

- Mohammad Asjad

[Salesforce AI Research Proposes a Novel Threat Model: Building Secure LLM Applications Against Prompt Leakage Attacks](/content/2024/10/04/salesforce-ai-research-proposes-a-novel-threat-model-building-secure-llm-applications-against-prompt-leakage-attacks/index.html)

- Mohammad Asjad

[Evaluating the Vulnerabilities of Unlearning Techniques in Large Language Models: A Comprehensive White-Box Analysis](/content/2024/10/03/evaluating-the-vulnerabilities-of-unlearning-techniques-in-large-language-models-a-comprehensive-white-box-analysis/index.html)

- Mohammad Asjad

[Logic-of-Thought: Enhancing Logical Reasoning in Large Language Models through Propositional Logic Augmentation](/content/2024/10/02/logic-of-thought-enhancing-logical-reasoning-in-large-language-models-through-propositional-logic-augmentation/index.html)

- Mohammad Asjad

[Model Collapse in the Synthetic Data Era: Analytical Insights and Mitigation Strategies](/content/2024/10/01/model-collapse-in-the-synthetic-data-era-analytical-insights-and-mitigation-strategies/index.html)

- Mohammad Asjad

[WaveletGPT: Leveraging Wavelet Theory for Speedier LLM Training Across Modalities](/content/2024/09/30/waveletgpt-leveraging-wavelet-theory-for-speedier-llm-training-across-modalities/index.html)

- Mohammad Asjad

[Scaling Laws and Model Comparison: New Frontiers in Large-Scale Machine Learning](/content/2024/09/29/scaling-laws-and-model-comparison-new-frontiers-in-large-scale-machine-learning/index.html)

- Mohammad Asjad

[RxEnvironments.jl: A Reactive Programming Approach to Complex Agent-Environment Simulations in the Julia Language](/content/2024/09/27/rxenvironments-jl-a-reactive-programming-approach-to-complex-agent-environment-simulations-in-the-julia-language/index.html)

- Mohammad Asjad

[Bridging Policy and Practice: Transparency Reporting in Foundation Models](/content/2024/09/27/bridging-policy-and-practice-transparency-reporting-in-foundation-models/index.html)

- Mohammad Asjad

[Researchers from John Hopkins and Samaya AI Propose Promptriever: A Zero-Shot Promptable Retriever Trained from a New Instruction-based Retrieval Dataset](/content/2024/09/26/researchers-from-john-hopkins-and-samaya-ai-propose-promptriever-a-zero-shot-promptable-retriever-trained-from-a-new-instruction-based-retrieval-dataset/index.html)

- Mohammad Asjad

[Iteration of Thought: An AI Framework for Enhancing LLM Responses by Generating “thought”-Provoking Prompts](/content/2024/09/25/iteration-of-thought-an-ai-framework-for-enhancing-llm-responses-by-generating-thought-provoking-prompts/index.html)

- Mohammad Asjad

[RetrievalAttention: A Training-Free Machine Learning Approach to both Accelerate Attention Computation and Reduce GPU Memory Consumption](/content/2024/09/24/retrievalattention-a-training-free-machine-learning-approach-to-both-accelerate-attention-computation-and-reduce-gpu-memory-consumption/index.html)

- Mohammad Asjad

[DCMAC: Demand-Aware Customized Communication for Efficient Multi-Agent Reinforcement Learning](/content/2024/09/23/dcmac-demand-aware-customized-communication-for-efficient-multi-agent-reinforcement-learning/index.html)

- Mohammad Asjad

[CORE-Bench: A Benchmark Consisting of 270 Tasks based on 90 Scientific Papers Across Computer Science, Social Science, and Medicine with Python or R Codebases](/content/2024/09/22/core-bench-a-benchmark-consisting-of-270-tasks-based-on-90-scientific-papers-across-computer-science-social-science-and-medicine-with-python-or-r-codebases/index.html)

- Mohammad Asjad

[Gated Slot Attention: Advancing Linear Attention Models for Efficient and Effective Language Processing](/content/2024/09/21/gated-slot-attention-advancing-linear-attention-models-for-efficient-and-effective-language-processing/index.html)

- Mohammad Asjad

[Sketch: An Innovative AI Toolkit Designed to Streamline LLM Operations Across Diverse Fields](/content/2024/09/20/sketch-an-innovative-ai-toolkit-designed-to-streamline-llm-operations-across-diverse-fields/index.html)

- Mohammad Asjad

[This AI Paper from Centre for the Governance of AI Proposes a Grading Rubric for AI Safety Frameworks](/content/2024/09/19/this-ai-paper-from-centre-for-the-governance-of-ai-proposes-a-grading-rubric-for-ai-safety-frameworks/index.html)

- Mohammad Asjad

[Contrastive Twist Learning and Bidirectional SMC Bounds: A New Paradigm for Language Model Control](/content/2024/09/18/contrastive-twist-learning-and-bidirectional-smc-bounds-a-new-paradigm-for-language-model-control/index.html)

- Mohammad Asjad

[Rethinking LLM Training: The Promise of Inverse Reinforcement Learning Techniques](/content/2024/09/16/rethinking-llm-training-the-promise-of-inverse-reinforcement-learning-techniques/index.html)

- Mohammad Asjad

[Google DeepMind Researchers Propose Human-Centric Alignment for Vision Models to Boost AI Generalization and Interpretation](/content/2024/09/16/google-deepmind-researchers-propose-human-centric-alignment-for-vision-models-to-boost-ai-generalization-and-interpretation/index.html)

- Mohammad Asjad

[LLaMA-Omni: A Novel AI Model Architecture Designed for Low-Latency and High-Quality Speech Interaction with LLMs](/content/2024/09/15/llama-omni-a-novel-ai-model-architecture-designed-for-low-latency-and-high-quality-speech-interaction-with-llms/index.html)

- Mohammad Asjad

[Small but Mighty: The Enduring Relevance of Small Language Models in the Age of LLMs](/content/2024/09/15/small-but-mighty-the-enduring-relevance-of-small-language-models-in-the-age-of-llms/index.html)

- Mohammad Asjad

[Ebay Researchers Introduce GraphEx: A Graph-based Extraction Method for Advertiser Keyphrase Recommendation](/content/2024/09/14/ebay-researchers-introduce-graphex-a-graph-based-extraction-method-for-advertiser-keyphrase-recommendation/index.html)

- Mohammad Asjad

[Automating Reinforcement Learning Workflows with Vision-Language Models: Towards Autonomous Mastery of Robotic Tasks](/content/2024/09/13/automating-reinforcement-learning-workflows-with-vision-language-models-towards-autonomous-mastery-of-robotic-tasks/index.html)

- Mohammad Asjad

[FlashSigmoid: A Hardware-Aware and Memory-Efficient Implementation of Sigmoid Attention Yielding a 17% Inference Kernel Speed-Up over FlashAttention-2 on H100 GPUs](/content/2024/09/13/flashsigmoid-a-hardware-aware-and-memory-efficient-implementation-of-sigmoid-attention-yielding-a-17-inference-kernel-speed-up-over-flashattention-2-on-h100-gpus/index.html)

- Mohammad Asjad

[Apple Researchers Propose a Novel AI Algorithm to Optimize a Byte-Level Representation for Automatic Speech Recognition ASR and Compare it with UTF-8 Representation](/content/2024/09/11/apple-researchers-propose-a-novel-ai-algorithm-to-optimize-a-byte-level-representation-for-automatic-speech-recognition-asr-and-compare-it-with-utf-8-representation/index.html)

- Mohammad Asjad

[Optimizing Document Understanding with DocOwl2: A Novel High-Resolution Compression Architecture](/content/2024/09/11/optimizing-document-understanding-with-docowl2-a-novel-high-resolution-compression-architecture/index.html)

- Mohammad Asjad

[Language-Guided World Models (LWMs): Enhancing Agent Controllability and Compositional Generalization through Natural Language](/content/2024/09/11/language-guided-world-models-lwms-enhancing-agent-controllability-and-compositional-generalization-through-natural-language/index.html)

- Mohammad Asjad

[Political DEBATE Language Models: Open-Source Solutions for Efficient Text Classification in Political Science](/content/2024/09/09/political-debate-language-models-open-source-solutions-for-efficient-text-classification-in-political-science/index.html)

- Mohammad Asjad

[VQ4DiT: A Fast Post-Training Vector Quantization Method for DiTs (Diffusion Transformers Models)](/content/2024/09/09/vq4dit-a-fast-post-training-vector-quantization-method-for-dits-diffusion-transformers-models/index.html)

- Mohammad Asjad

[Advancing Cantonese NLP: Bridging Development Gaps in Large Language Models with New Benchmarks and Open-Source Innovations](/content/2024/09/08/advancing-cantonese-nlp-bridging-development-gaps-in-large-language-models-with-new-benchmarks-and-open-source-innovations/index.html)

- Mohammad Asjad

[OpenFGL: A Comprehensive Benchmark for Advancing Federated Graph Learning](/content/2024/09/07/openfgl-a-comprehensive-benchmark-for-advancing-federated-graph-learning/index.html)

- Mohammad Asjad

[SFR-GNN: A Novel Graph Neural Networks (GNN) Model that Employs an ‘Attribute Pre-Training and Structure Fine-Tuning’ Strategy to Achieve Robustness Against Structural Attacks](/content/2024/09/07/sfr-gnn-a-novel-graph-neural-networks-gnn-model-that-employs-an-attribute-pre-training-and-structure-fine-tuning-strategy-to-achieve-robustness-against-structural-attacks/index.html)

- Mohammad Asjad

[Comparative Analysis of LLM and Traditional Text Augmentation: Accuracy, Efficiency, and Cost-Effectiveness](/content/2024/09/06/comparative-analysis-of-llm-and-traditional-text-augmentation-accuracy-efficiency-and-cost-effectiveness/index.html)

- Mohammad Asjad

[Why GPU Utilization Falls Short: Understanding Streaming Multiprocessor (SM) Efficiency for Better LLM Performance](/content/2024/09/03/why-gpu-utilization-falls-short-understanding-streaming-multiprocessor-sm-efficiency-for-better-llm-performance/index.html)

- Mohammad Asjad

[LLaVaOLMoBitnet1B: The First Ternary Multimodal LLM Capable of Accepting Image(s) and Text Inputs to Produce Coherent Textual Response](/content/2024/09/03/llavaolmobitnet1b-the-first-ternary-multimodal-llm-capable-of-accepting-images-and-text-inputs-to-produce-coherent-textual-response/index.html)

- Mohammad Asjad

[The Mamba in the Llama: Accelerating Inference with Speculative Decoding](/content/2024/09/01/the-mamba-in-the-llama-accelerating-inference-with-speculative-decoding/index.html)

- Mohammad Asjad

[Qwen2-VL Released: The Latest Version of the Vision Language Models based on Qwen2 in the Qwen Model Familities](/content/2024/09/01/qwen2-vl-released-the-latest-version-of-the-vision-language-models-based-on-qwen2-in-the-qwen-model-familities/index.html)

- Mohammad Asjad

[Microsoft Researchers Combine Small and Large Language Models for Faster, More Accurate Hallucination Detection](/content/2024/08/31/microsoft-researchers-combine-small-and-large-language-models-for-faster-more-accurate-hallucination-detection/index.html)

- Mohammad Asjad

[Aleph Alpha Researchers Release Pharia-1-LLM-7B: Two Distinct Variants- Pharia-1-LLM-7B-Control and Pharia-1-LLM-7B-Control-Aligned](/content/2024/08/30/aleph-alpha-researchers-release-pharia-1-llm-7b-two-distinct-variants-pharia-1-llm-7b-control-and-pharia-1-llm-7b-control-aligned/index.html)

- Mohammad Asjad

[Jina AI Introduced ‘Late Chunking’: A Simple AI Approach to Embed Short Chunks by Leveraging the Power of Long-Context Embedding Models](/content/2024/08/27/jina-ai-introduced-late-chunking-a-simple-ai-approach-to-embed-short-chunks-by-leveraging-the-power-of-long-context-embedding-models/index.html)

- Mohammad Asjad

[Humboldt: A Specification-based System Framework for Generating a Data Discovery UI from Different Metadata Providers](/content/2024/08/26/humboldt-a-specification-based-system-framework-for-generating-a-data-discovery-ui-from-different-metadata-providers/index.html)

- Mohammad Asjad

[Improving RLHF (Reinforcement Learning from Human Feedback) with Critique-Generated Reward Models](/content/2024/08/25/improving-rlhf-reinforcement-learning-from-human-feedback-with-critique-generated-reward-models/index.html)

- Mohammad Asjad

[AWS Enhancing Information Retrieval in Large Language Models: A Data-Centric Approach Using Metadata, Synthetic QAs, and Meta Knowledge Summaries for Improved Accuracy and Relevancy](/content/2024/08/24/aws-enhancing-information-retrieval-in-large-language-models-a-data-centric-approach-using-metadata-synthetic-qas-and-meta-knowledge-summaries-for-improved-accuracy-and-relevancy/index.html)

- Mohammad Asjad

[Integrating Graph Structures into Language Models: A Comprehensive Study of GraphRAG](/content/2024/08/24/integrating-graph-structures-into-language-models-a-comprehensive-study-of-graphrag/index.html)

- Mohammad Asjad

[Code as a Catalyst: Improving LLM Capabilities Across Diverse Tasks](/content/2024/08/22/code-as-a-catalyst-improving-llm-capabilities-across-diverse-tasks/index.html)

- Mohammad Asjad

[DaRec: A Novel Plug-and-Play Alignment Framework for LLMs and Collaborative Models](/content/2024/08/21/darec-a-novel-plug-and-play-alignment-framework-for-llms-and-collaborative-models/index.html)

- Mohammad Asjad

[DataVisT5: A Powerful Pre-Trained Language Model for Seamless Data Visualization Tasks](/content/2024/08/20/datavist5-a-powerful-pre-trained-language-model-for-seamless-data-visualization-tasks/index.html)

- Mohammad Asjad

[Improving Robustness Against Bias in Social Science Machine Learning: The Promise of Instruction-Based Models](/content/2024/08/19/improving-robustness-against-bias-in-social-science-machine-learning-the-promise-of-instruction-based-models/index.html)

- Mohammad Asjad

[UniBench: A Python Library to Evaluate Vision-Language Models VLMs Robustness Across Diverse Benchmarks](/content/2024/08/18/unibench-a-python-library-to-evaluate-vision-language-models-vlms-robustness-across-diverse-benchmarks/index.html)

- Mohammad Asjad

[Meta AI and NYU Researchers Propose E-RLHF to Combat LLM Jailbreaking](/content/2024/08/18/meta-ai-and-nyu-researchers-propose-e-rlhf-to-combat-llm-jailbreaking/index.html)

- Mohammad Asjad

[Google AI Announces Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](/content/2024/08/17/google-ai-announces-scaling-llm-test-time-compute-optimally-can-be-more-effective-than-scaling-model-parameters/index.html)

- Mohammad Asjad

[DeepSeek-AI Open-Sources DeepSeek-Prover-V1.5: A Language Model with 7 Billion Parameters that Outperforms all Open-Source Models in Formal Theorem Proving in Lean 4](/content/2024/08/17/deepseek-ai-open-sources-deepseek-prover-v1-5-a-language-model-with-7-billion-parameters-that-outperforms-all-open-source-models-in-formal-theorem-proving-in-lean-4/index.html)

- Mohammad Asjad

[Answer.AI Releases answerai-colbert-small: A Proof of Concept for Smaller, Faster, Modern ColBERT Models](/content/2024/08/16/answer-ai-releases-answerai-colbert-small-a-proof-of-concept-for-smaller-faster-modern-colbert-models/index.html)

- Mohammad Asjad

[Self-play muTuAl Reasoning (rStar): A Novel AI Approach that Boosts Small Language Models SLMs’ Reasoning Capability during Inference without Fine-Tuning](/content/2024/08/13/self-play-mutual-reasoning-rstar-a-novel-ai-approach-that-boosts-small-language-models-slms-reasoning-capability-during-inference-without-fine-tuning/index.html)

- Mohammad Asjad

[NACL: A Robust KV Cache Eviction Framework for Efficient Long-Text Processing in LLMs](/content/2024/08/12/nacl-a-robust-kv-cache-eviction-framework-for-efficient-long-text-processing-in-llms/index.html)

- Mohammad Asjad

[Apple Researchers Present KGLens: A Novel AI Method Tailored for Visualizing and Evaluating the Factual Knowledge Embedded in LLMs](/content/2024/08/12/apple-researchers-present-kglens-a-novel-ai-method-tailored-for-visualizing-and-evaluating-the-factual-knowledge-embedded-in-llms/index.html)

- Mohammad Asjad

[Parler-TTS Released: A Fully Open-Sourced Text-to-Speech Model with Advanced Speech Synthesis for Complex and Lightweight Applications](/content/2024/08/10/parler-tts-released-a-fully-open-sourced-text-to-speech-model-with-advanced-speech-synthesis-for-complex-and-lightweight-applications/index.html)

- Mohammad Asjad

[RadGraph2: A New Dataset for Tracking Disease Progression in Radiology Reports](/content/2024/08/10/radgraph2-a-new-dataset-for-tracking-disease-progression-in-radiology-reports/index.html)

- Mohammad Asjad

[NYU Researchers Open-Sourced GPUDrive: A GPU-Accelerated Multi-Agent Driving Simulation at 1 Million FPS](/content/2024/08/08/nyu-researchers-open-sourced-gpudrive-a-gpu-accelerated-multi-agent-driving-simulation-at-1-million-fps/index.html)

- Mohammad Asjad

[Allen Institute for AI (AI2) Released a New Bundle of OLMo 1B and 7B Assets](/content/2024/08/06/allen-institute-for-ai-ai2-released-a-new-bundle-of-olmo-1b-and-7b-assets/index.html)

- Mohammad Asjad

[This AI Paper Introduces a Verbalized Way to Perform Machine Learning and Conducts Several Case Studies on Regression and Classification Tasks](/content/2024/08/05/this-ai-paper-introduces-a-verbalized-way-to-perform-machine-learning-and-conducts-several-case-studies-on-regression-and-classification-tasks/index.html)

- Mohammad Asjad

[Magpie-Ultra Dataset Released: Harnessing Llama 3.1 405B for Diverse AI Instruction-Response Pairs](/content/2024/08/04/magpie-ultra-dataset-released-harnessing-llama-3-1-405b-for-diverse-ai-instruction-response-pairs/index.html)

- Mohammad Asjad

[Character AI Releases Prompt Poet: A New Low Code Python Libary that Streamlines Prompt Design for both Developers and Non-Technical Users](/content/2024/08/03/character-ai-releases-prompt-poet-a-new-low-code-python-libary-that-streamlines-prompt-design-for-both-developers-and-non-technical-users/index.html)

- Mohammad Asjad

[MLPs vs KANs: Evaluating Performance in Machine Learning, Computer Vision, NLP, and Symbolic Tasks](/content/2024/08/03/mlps-vs-kans-evaluating-performance-in-machine-learning-computer-vision-nlp-and-symbolic-tasks/index.html)

- Mohammad Asjad

[Lyzr Automata: A Low-Code Multi-Agent Framework for Advanced Process Automation](/content/2024/08/03/lyzr-automata-a-low-code-multi-agent-framework-for-advanced-process-automation/index.html)

- Mohammad Asjad

[Black Forest Labs Open-Source FLUX.1: A 12 Billion Parameter Rectified Flow Transformer Capable of Generating Images from Text Descriptions](/content/2024/08/02/black-forest-labs-open-source-flux-1-a-12-billion-parameter-rectified-flow-transformer-capable-of-generating-images-from-text-descriptions/index.html)

- Mohammad Asjad

[Google AI Introduces ShieldGemma: A Comprehensive Suite of LLM-based Safety Content Moderation Models Built on Gemma2](/content/2024/08/02/google-ai-introduces-shieldgemma-a-comprehensive-suite-of-llm-based-safety-content-moderation-models-built-on-gemma2/index.html)

- Mohammad Asjad

[PersonaGym: A Dynamic AI Framework for Comprehensive Evaluation of LLM Persona Agents](/content/2024/08/02/personagym-a-dynamic-ai-framework-for-comprehensive-evaluation-of-llm-persona-agents/index.html)

- Mohammad Asjad

[Salesforce AI Introduces ‘ThinK’: A New AI Method that Exploits Substantial Redundancy Across the Channel Dimension of the KV Cache](/content/2024/08/01/salesforce-ai-introduces-think-a-new-ai-method-that-exploits-substantial-redundancy-across-the-channel-dimension-of-the-kv-cache/index.html)

- Mohammad Asjad

[rLLM (relationLLM): A PyTorch Library Designed for Relational Table Learning (RTL) with Large Language Models (LLMs)](/content/2024/07/30/rllm-relationllm-a-pytorch-library-designed-for-relational-table-learning-rtl-with-large-language-models-llms/index.html)

- Mohammad Asjad

[Recursive IntroSpEction (RISE): A Machine Learning Approach for Fine-Tuning LLMs to Improve Their Own Responses Over Multiple Turns Sequentially](/content/2024/07/29/recursive-introspection-rise-a-machine-learning-approach-for-fine-tuning-llms-to-improve-their-own-responses-over-multiple-turns-sequentially/index.html)

- Mohammad Asjad

[This AI Paper from Stanford Provides New Insights on AI Model Collapse and Data Accumulation](/content/2024/07/29/this-ai-paper-from-stanford-provides-new-insights-on-ai-model-collapse-and-data-accumulation/index.html)

- Mohammad Asjad

[CompeteAI: An Artificial Intelligence AI Framework that Understands the Competition Dynamics of Large Language Model-based Agents](/content/2024/07/27/competeai-an-artificial-intelligence-ai-framework-that-understands-the-competition-dynamics-of-large-language-model-based-agents/index.html)

- Mohammad Asjad

[FLUTE: A CUDA Kernel Designed for Fused Quantized Matrix Multiplications to Accelerate LLM Inference](/content/2024/07/26/flute-a-cuda-kernel-designed-for-fused-quantized-matrix-multiplications-to-accelerate-llm-inference/index.html)

- Mohammad Asjad

[SF-LLaVA: A Training-Free Video LLM that is Built Upon LLaVA-NeXT and Requires No Additional Fine-Tuning to Work Effectively for Various Video Tasks](/content/2024/07/25/sf-llava-a-training-free-video-llm-that-is-built-upon-llava-next-and-requires-no-additional-fine-tuning-to-work-effectively-for-various-video-tasks/index.html)

- Mohammad Asjad

[Apple Researchers Propose LazyLLM: A Novel AI Technique for Efficient LLM Inference in Particular under Long Context Scenarios](/content/2024/07/23/apple-researchers-propose-lazyllm-a-novel-ai-technique-for-efficient-llm-inference-in-particular-under-long-context-scenarios/index.html)

- Mohammad Asjad

[WTU-Eval: A New Standard Benchmark Tool for Evaluating Large Language Models LLMs Usage Capabilities](/content/2024/07/23/wtu-eval-a-new-standard-benchmark-tool-for-evaluating-large-language-models-llms-usage-capabilities/index.html)

- Mohammad Asjad

[Open Artificial Knowledge (OAK) Dataset: A Large-Scale Resource for AI Research Derived from Wikipedia’s Main Categories](/content/2024/07/22/open-artificial-knowledge-oak-dataset-a-large-scale-resource-for-ai-research-derived-from-wikipedias-main-categories/index.html)

- Mohammad Asjad

[From RAG to ReST: A Survey of Advanced Techniques in Large Language Model Development](/content/2024/07/22/from-rag-to-rest-a-survey-of-advanced-techniques-in-large-language-model-development/index.html)

- Mohammad Asjad

[Athene-Llama3-70B Released: An Open-Weight LLM Trained through RLHF based on Llama-3-70B-Instruct](/content/2024/07/21/athene-llama3-70b-released-an-open-weight-llm-trained-through-rlhf-based-on-llama-3-70b-instruct/index.html)

- Mohammad Asjad

[Agent Symbolic Learning: An Artificial Intelligence AI Framework for Agent Learning that Jointly Optimizes All Symbolic Components within an Agent System](/content/2024/07/21/agent-symbolic-learning-an-artificial-intelligence-ai-framework-for-agent-learning-that-jointly-optimizes-all-symbolic-components-within-an-agent-system/index.html)

- Mohammad Asjad

[ZebraLogic: A Logical Reasoning AI Benchmark Designed for Evaluating LLMs with Logic Puzzles](/content/2024/07/20/zebralogic-a-logical-reasoning-ai-benchmark-designed-for-evaluating-llms-with-logic-puzzles/index.html)

- Mohammad Asjad

[MUSE: A Comprehensive AI Framework for Evaluating Machine Unlearning in Language Models](/content/2024/07/20/muse-a-comprehensive-ai-framework-for-evaluating-machine-unlearning-in-language-models/index.html)

- Mohammad Asjad

[EM-LLM: A Novel and Flexible Architecture that Integrates Key Aspects of Human Episodic Memory and Event Cognition into Transformer-based Language Models](/content/2024/07/19/em-llm-a-novel-and-flexible-architecture-that-integrates-key-aspects-of-human-episodic-memory-and-event-cognition-into-transformer-based-language-models/index.html)

- Mohammad Asjad

[From Diagrams to Solutions: MAVIS’s Three-Stage Framework for Mathematical AI](/content/2024/07/19/from-diagrams-to-solutions-maviss-three-stage-framework-for-mathematical-ai/index.html)

- Mohammad Asjad

[DotaMath: Advancing LLMs’ Mathematical Reasoning Through Decomposition and Self-Correction](/content/2024/07/19/dotamath-advancing-llms-mathematical-reasoning-through-decomposition-and-self-correction/index.html)

- Mohammad Asjad

[G-Retriever: Advancing Real-World Graph Question Answering with RAG and LLMs](/content/2024/07/17/g-retriever-advancing-real-world-graph-question-answering-with-rag-and-llms/index.html)

- Mohammad Asjad

[MELLE: A Novel Continuous-Valued Tokens-based Language Modeling Approach for Text-to-Speech Synthesis (TTS)](/content/2024/07/17/melle-a-novel-continuous-valued-tokens-based-language-modeling-approach-for-text-to-speech-synthesis-tts/index.html)

- Mohammad Asjad

[UCSD Researchers Propose a General Variational Inference-based Framework (MCD) to Infer the Underlying Causal Models as well as the Mixing Probability of Each Sample](/content/2024/07/17/ucsd-researchers-propose-a-general-variational-inference-based-framework-mcd-to-infer-the-underlying-causal-models-as-well-as-the-mixing-probability-of-each-sample/index.html)

- Mohammad Asjad

[Planetarium: A New Benchmark to Evaluate LLMs on Translating Natural Language Descriptions of Planning Problems into Planning Domain Definition Language PDDL](/content/2024/07/15/planetarium-a-new-benchmark-to-evaluate-llms-on-translating-natural-language-descriptions-of-planning-problems-into-planning-domain-definition-language-pddl/index.html)

- Mohammad Asjad

[Samsung Researchers Introduce LoRA-Guard: A Parameter-Efficient Guardrail Adaptation Method that Relies on Knowledge Sharing between LLMs and Guardrail Models](/content/2024/07/14/samsung-researchers-introduce-lora-guard-a-parameter-efficient-guardrail-adaptation-method-that-relies-on-knowledge-sharing-between-llms-and-guardrail-models/index.html)

- Mohammad Asjad

[InternLM-XComposer-2.5 (IXC-2.5): A Versatile Large-Vision Language Model that Supports Long-Contextual Input and Output](/content/2024/07/13/internlm-xcomposer-2-5-ixc-2-5-a-versatile-large-vision-language-model-that-supports-long-contextual-input-and-output/index.html)

- Mohammad Asjad

[Can LLMs Help Accelerate the Discovery of Data-Driven Scientific Hypotheses? Meet DiscoveryBench: A Comprehensive LLM Benchmark that Formalizes the Multi-Step Process of Data-Driven Discovery](/content/2024/07/13/can-llms-help-accelerate-the-discovery-of-data-driven-scientific-hypotheses-meet-discoverybench-a-comprehensive-llm-benchmark-that-formalizes-the-multi-step-process-of-data-driven-discovery/index.html)

- Mohammad Asjad

[Google DeepMind Unveils PaliGemma: A Versatile 3B Vision-Language Model VLM with Large-Scale Ambitions](/content/2024/07/12/google-deepmind-unveils-paligemma-a-versatile-3b-vision-language-model-vlm-with-large-scale-ambitions/index.html)

- Mohammad Asjad

[This AI Paper from the National University of Singapore Introduces a Defense Against Adversarial Attacks on LLMs Utilizing Self-Evaluation](/content/2024/07/10/this-ai-paper-from-the-national-university-of-singapore-introduces-a-defense-against-adversarial-attacks-on-llms-utilizing-self-evaluation/index.html)

- Mohammad Asjad

[NVIDIA Introduces RankRAG: A Novel RAG Framework that Instruction-Tunes a Single LLM for the Dual Purposes of Top-k Context Ranking and Answer Generation in RAG](/content/2024/07/09/nvidia-introduces-rankrag-a-novel-rag-framework-that-instruction-tunes-a-single-llm-for-the-dual-purposes-of-top-k-context-ranking-and-answer-generation-in-rag/index.html)

- Mohammad Asjad

[This AI Research from Tenyx Explore the Reasoning Abilities of Large Language Models (LLMs) Through Their Geometrical Understanding](/content/2024/07/08/this-ai-research-from-tenyx-explore-the-reasoning-abilities-of-large-language-models-llms-through-their-geometrical-understanding/index.html)

- Mohammad Asjad

[WorldBench: A Dynamic and Flexible LLM Benchmark Composed of Per-Country Data from the World Bank](/content/2024/07/07/worldbench-a-dynamic-and-flexible-llm-benchmark-composed-of-per-country-data-from-the-world-bank/index.html)

- Mohammad Asjad

[Exploring the Influence of AI-Based Recommenders on Human Behavior: Methodologies, Outcomes, and Future Research Directions](/content/2024/07/06/exploring-the-influence-of-ai-based-recommenders-on-human-behavior-methodologies-outcomes-and-future-research-directions/index.html)

- Mohammad Asjad

[Safeguarding Healthcare AI: Exposing and Addressing LLM Manipulation Risks](/content/2024/07/06/safeguarding-healthcare-ai-exposing-and-addressing-llm-manipulation-risks/index.html)

- Mohammad Asjad

[A Concurrent Programming Framework for Quantitative Analysis of Efficiency Issues When Serving Multiple Long-Context Requests Under Limited GPU High-Bandwidth Memory (HBM) Regime](/content/2024/07/05/a-concurrent-programming-framework-for-quantitative-analysis-of-efficiency-issues-when-serving-multiple-long-context-requests-under-limited-gpu-high-bandwidth-memory-hbm-regime/index.html)

- Mohammad Asjad

[Rethinking QA Dataset Design: How Popular Knowledge Enhances LLM Accuracy?](/content/2024/07/04/rethinking-qa-dataset-design-how-popular-knowledge-enhances-llm-accuracy/index.html)

- Mohammad Asjad

[EvoAgent: A Generic Method to Automatically Extend Expert Agents to Multi-Agent Systems via the Evolutionary Algorithm](/content/2024/07/03/evoagent-a-generic-method-to-automatically-extend-expert-agents-to-multi-agent-systems-via-the-evolutionary-algorithm/index.html)

- Mohammad Asjad

[Privacy Meets Performance: GPT4All 3.0 Redefines Local AI Interaction](/content/2024/07/03/privacy-meets-performance-gpt4all-3-0-redefines-local-ai-interaction/index.html)

- Mohammad Asjad

[45 Shades of AI Safety: SORRY-Bench’s Innovative Taxonomy for LLM Refusal Behavior Analysis](/content/2024/07/02/45-shades-of-ai-safety-sorry-benchs-innovative-taxonomy-for-llm-refusal-behavior-analysis/index.html)

- Mohammad Asjad

[Researchers at Princeton University Proposes Edge Pruning: An Effective and Scalable Method for Automated Circuit Finding](/content/2024/07/02/researchers-at-princeton-university-proposes-edge-pruning-an-effective-and-scalable-method-for-automated-circuit-finding/index.html)

- Mohammad Asjad

[Fal AI Introduces AuraSR: A 600M Parameter Upsampler Model Derived from the GigaGAN](/content/2024/07/01/fal-ai-introduces-aurasr-a-600m-parameter-upsampler-model-derived-from-the-gigagan/index.html)

- Mohammad Asjad

[Researchers at Brown University Explore Zero-Shot Cross-Lingual Generalization of Preference Tuning in Detoxifying LLMs](/content/2024/06/30/researchers-at-brown-university-explore-zero-shot-cross-lingual-generalization-of-preference-tuning-in-detoxifying-llms/index.html)

- Mohammad Asjad

[This AI Paper from CMU and Google DeepMind Studies the Role of Synthetic Data for Improving Math Reasoning Capabilities of LLMs](/content/2024/06/30/this-ai-paper-from-cmu-and-google-deepmind-studies-the-role-of-synthetic-data-for-improving-math-reasoning-capabilities-of-llms/index.html)

- Mohammad Asjad

[MuxServe: A Flexible and Efficient Spatial-Temporal Multiplexing System to Serve Multiple LLMs Concurrently](/content/2024/06/30/muxserve-a-flexible-and-efficient-spatial-temporal-multiplexing-system-to-serve-multiple-llms-concurrently/index.html)

- Mohammad Asjad

[Q\*: A Versatile Artificial Intelligence AI Approach to Improve LLM Performance in Reasoning Tasks](/content/2024/06/27/q-a-versatile-artificial-intelligence-ai-approach-to-improve-llm-performance-in-reasoning-tasks/index.html)

- Mohammad Asjad

[GraphReader: A Graph-based AI Agent System Designed to Handle Long Texts by Structuring them into a Graph and Employing an Agent to Explore this Graph Autonomously](/content/2024/06/26/graphreader-a-graph-based-ai-agent-system-designed-to-handle-long-texts-by-structuring-them-into-a-graph-and-employing-an-agent-to-explore-this-graph-autonomously/index.html)

- Mohammad Asjad

[Camb AI Releases MARS5 TTS: A Novel Open Source Text to Speech Model for Insane Prosody](/content/2024/06/26/camb-ai-releases-mars5-tts-a-novel-open-source-text-to-speech-model-for-insane-prosody/index.html)

- Mohammad Asjad

[Whiteboard-of-Thought (WoT) Prompting: A Simple AI Approach to Enhance the Visual Reasoning Abilities of MLLMs Across Modalities](/content/2024/06/24/whiteboard-of-thought-wot-prompting-a-simple-ai-approach-to-enhance-the-visual-reasoning-abilities-of-mllms-across-modalities/index.html)

- Mohammad Asjad

[MIPRO: A Novel Optimizer that Outperforms Baselines on Five of Six Diverse Language Model LM Programs Using a Best-in-Class Open-Source Model (Llama-3-8B) by 12.9% accuracy](/content/2024/06/24/mipro-a-novel-optimizer-that-outperforms-baselines-on-five-of-six-diverse-language-model-lm-programs-using-a-best-in-class-open-source-model-llama-3-8b-by-12-9-accuracy/index.html)

- Mohammad Asjad

[Cephalo: A Series of Open-Source Multimodal Vision Large Language Models (V-LLMs) Specifically in the Context of Bio-Inspired Design](/content/2024/06/23/cephalo-a-series-of-open-source-multimodal-vision-large-language-models-v-llms-specifically-in-the-context-of-bio-inspired-design/index.html)

- Mohammad Asjad

[LOFT: A Comprehensive AI Benchmark for Evaluating Long-Context Language Models](/content/2024/06/23/loft-a-comprehensive-ai-benchmark-for-evaluating-long-context-language-models/index.html)

- Mohammad Asjad

[The Rise of Diffusion-Based Language Models: Comparing SEDD and GPT-2](/content/2024/06/22/the-rise-of-diffusion-based-language-models-comparing-sedd-and-gpt-2/index.html)

- Mohammad Asjad

[PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers](/content/2024/06/22/planrag-a-plan-then-retrieval-augmented-generation-for-generative-large-language-models-as-decision-makers/index.html)

- Mohammad Asjad

[CS-Bench: A Bilingual (Chinese-English) Benchmark Dedicated to Evaluating the Performance of LLMs in Computer Science](/content/2024/06/20/cs-bench-a-bilingual-chinese-english-benchmark-dedicated-to-evaluating-the-performance-of-llms-in-computer-science/index.html)

- Mohammad Asjad

[StreamSpeech: A Direct Simul-S2ST Speech-to-Speech Translation Model that Jointly Learns Translation and Simultaneous Policy in a Unified Framework of Multi-Task Learning](/content/2024/06/20/streamspeech-a-direct-simul-s2st-speech-to-speech-translation-model-that-jointly-learns-translation-and-simultaneous-policy-in-a-unified-framework-of-multi-task-learning/index.html)

- Mohammad Asjad

[Apple Releases 4M-21: A Very Effective Multimodal AI Model that Solves Tens of Tasks and Modalities](/content/2024/06/18/apple-releases-4m-21-a-very-effective-multimodal-ai-model-that-solves-tens-of-tasks-and-modalities/index.html)

- Mohammad Asjad

[Pixel Transformer: Challenging Locality Bias in Vision Models](/content/2024/06/17/pixel-transformer-challenging-locality-bias-in-vision-models/index.html)

- Mohammad Asjad

[Neural Algorithmic Reasoning for Transformers: The TransNAR Framework](/content/2024/06/16/neural-algorithmic-reasoning-for-transformers-the-transnar-framework/index.html)

- Mohammad Asjad

[Microsoft Researchers Introduce Samba 3.8B: A Simple Mamba+Sliding Window Attention Architecture that Outperforms Phi3-mini on Major Benchmarks](/content/2024/06/15/microsoft-researchers-introduce-samba-3-8b-a-simple-mambasliding-window-attention-architecture-that-outperforms-phi3-mini-on-major-benchmarks/index.html)

- Mohammad Asjad

[Enhancing Trust in Large Language Models: Fine-Tuning for Calibrated Uncertainties in High-Stakes Applications](/content/2024/06/15/enhancing-trust-in-large-language-models-fine-tuning-for-calibrated-uncertainties-in-high-stakes-applications/index.html)

- Mohammad Asjad

[SelfGoal: An Artificial Intelligence AI Framework to Enhance an LLM-based Agent’s Capabilities to Achieve High-Level Goals](/content/2024/06/14/selfgoal-an-artificial-intelligence-ai-framework-to-enhance-an-llm-based-agents-capabilities-to-achieve-high-level-goals/index.html)

- Mohammad Asjad

[GenAI-Arena: An Open Platform for Community-Based Evaluation of Generative AI Models](/content/2024/06/12/genai-arena-an-open-platform-for-community-based-evaluation-of-generative-ai-models/index.html)

- Mohammad Asjad

[Benchmarking Federated Learning for Large Language Models with FedLLM-Bench](/content/2024/06/11/benchmarking-federated-learning-for-large-language-models-with-fedllm-bench/index.html)

- Mohammad Asjad

[Advancing Reliable Question Answering with the CRAG Benchmark](/content/2024/06/11/advancing-reliable-question-answering-with-the-crag-benchmark/index.html)

- Mohammad Asjad

[From Low-Level to High-Level Tasks: Scaling Fine-Tuning with the ANDROIDCONTROL Dataset](/content/2024/06/10/from-low-level-to-high-level-tasks-scaling-fine-tuning-with-the-androidcontrol-dataset/index.html)

- Mohammad Asjad

[The Missing Piece: Combining Foundation Models and Open-Endedness for Artificial Superhuman Intelligence ASI](/content/2024/06/08/the-missing-piece-combining-foundation-models-and-open-endedness-for-artificial-superhuman-intelligence-asi/index.html)

- Mohammad Asjad

[Researchers at UC Berkeley Propose a Neural Diffusion Model that Operates on Syntax Trees for Program Synthesis](/content/2024/06/07/researchers-at-uc-berkeley-propose-a-neural-diffusion-model-that-operates-on-syntax-trees-for-program-synthesis/index.html)

- Mohammad Asjad

[Modeling Cultural Accumulation in Artificial Reinforcement Learning Agents](/content/2024/06/07/modeling-cultural-accumulation-in-artificial-reinforcement-learning-agents/index.html)

- Mohammad Asjad

[Quantized Eigenvector Matrices for 4-bit Second-Order Optimization of Deep Neural Networks](/content/2024/06/06/quantized-eigenvector-matrices-for-4-bit-second-order-optimization-of-deep-neural-networks/index.html)

- Mohammad Asjad

[Meet Tsinghua University’s GLM-4-9B-Chat-1M: An Outstanding Language Model Challenging GPT 4V, Gemini Pro (on vision), Mistral and Llama 3 8B](/content/2024/06/05/meet-tsinghua-universitys-glm-4-9b-chat-1m-an-outstanding-language-model-challenging-gpt-4v-gemini-pro-on-vision-mistral-and-llama-3-8b/index.html)

- Mohammad Asjad

[Parrot: Optimizing End-to-End Performance in LLM Applications Through Semantic Variables](/content/2024/06/03/parrot-optimizing-end-to-end-performance-in-llm-applications-through-semantic-variables/index.html)

- Mohammad Asjad

[Researchers at Microsoft Introduce Aurora: A Large-Scale Foundation Model of the Atmosphere Trained on Over a Million Hours of Diverse Weather and Climate Data](/content/2024/06/03/researchers-at-microsoft-introduce-aurora-a-large-scale-foundation-model-of-the-atmosphere-trained-on-over-a-million-hours-of-diverse-weather-and-climate-data/index.html)

- Mohammad Asjad

[Contextual Position Encoding (CoPE): A New Position Encoding Method that Allows Positions to be Conditioned on Context by Incrementing Position only on Certain Tokens Determined by the Model](/content/2024/06/02/contextual-position-encoding-cope-a-new-position-encoding-method-that-allows-positions-to-be-conditioned-on-context-by-incrementing-position-only-on-certain-tokens-determined-by-the-model/index.html)

- Mohammad Asjad

[GNN-RAG: A Novel AI Method for Combining Language Understanding Abilities of LLMs with the Reasoning Abilities of GNNs in a Retrieval-Augmented Generation (RAG) Style](/content/2024/06/01/gnn-rag-a-novel-ai-method-for-combining-language-understanding-abilities-of-llms-with-the-reasoning-abilities-of-gnns-in-a-retrieval-augmented-generation-rag-style/index.html)

- Mohammad Asjad

[This AI Paper from Princeton and the University of Warwick Proposes a Novel Artificial Intelligence Approach to Enhance the Utility of LLMs as Cognitive Models](/content/2024/06/01/this-ai-paper-from-princeton-and-the-university-of-warwick-proposes-a-novel-artificial-intelligence-approach-to-enhance-the-utility-of-llms-as-cognitive-models/index.html)

- Mohammad Asjad

[MoEUT: A Robust Machine Learning Approach to Addressing Universal Transformers’ Efficiency Challenges](/content/2024/05/31/moeut-a-robust-machine-learning-approach-to-addressing-universal-transformers-efficiency-challenges/index.html)

- Mohammad Asjad

[Llama3-V: A SOTA Open-Source VLM Model Comparable performance to GPT4-V, Gemini Ultra, Claude Opus with a 100x Smaller Model](/content/2024/05/31/llama3-v-a-sota-open-source-vlm-model-comparable-performance-to-gpt4-v-gemini-ultra-claude-opus-with-a-100x-smaller-model/index.html)

- Mohammad Asjad

[In-Context Learning Capabilities of Multi-Layer Perceptrons MLPs: A Comparative Study with Transformers](/content/2024/05/30/in-context-learning-capabilities-of-multi-layer-perceptrons-mlps-a-comparative-study-with-transformers/index.html)

- Mohammad Asjad

[Question-Answer Cross Attention Networks (QAN): Advancing Answer Selection in Community Question Answering](/content/2024/05/29/question-answer-cross-attention-networks-qan-advancing-answer-selection-in-community-question-answering/index.html)

- Mohammad Asjad

[Inductive Biases in Deep Learning: Understanding Feature Representation](/content/2024/05/28/inductive-biases-in-deep-learning-understanding-feature-representation/index.html)

- Mohammad Asjad

[Optimizing Agent Planning: A Parametric AI Approach to World Knowledge](/content/2024/05/27/optimizing-agent-planning-a-parametric-ai-approach-to-world-knowledge/index.html)

- Mohammad Asjad

[Unlocking the Potential of SirLLM: Advancements in Memory Retention and Attention Mechanisms](/content/2024/05/27/unlocking-the-potential-of-sirllm-advancements-in-memory-retention-and-attention-mechanisms/index.html)

- Mohammad Asjad

[Achieving Balance in Lifelong Learning: The WISE Memory Approach](/content/2024/05/26/achieving-balance-in-lifelong-learning-the-wise-memory-approach/index.html)

- Mohammad Asjad

[A Paradigm Shift: MoRA’s Role in Advancing Parameter-Efficient Fine-Tuning Techniques](/content/2024/05/25/a-paradigm-shift-moras-role-in-advancing-parameter-efficient-fine-tuning-techniques/index.html)

- Mohammad Asjad

[Transparency in Foundation Models: The Next Step in Foundation Model Transparency Index FMTI](/content/2024/05/25/transparency-in-foundation-models-the-next-step-in-foundation-model-transparency-index-fmti/index.html)

- Mohammad Asjad

[An Efficient AI Approach to Memory Reduction and Throughput Enhancement in LLMs](/content/2024/05/23/an-efficient-ai-approach-to-memory-reduction-and-throughput-enhancement-in-llms/index.html)

- Mohammad Asjad

[Apple Researchers Propose KV-Runahead: An Efficient Parallel LLM Inference Technique to Minimize the Time-to-First-Token](/content/2024/05/22/apple-researchers-propose-kv-runahead-an-efficient-parallel-llm-inference-technique-to-minimize-the-time-to-first-token/index.html)

- Mohammad Asjad

[Toward Responsible Innovation: Evaluating Risks and Opportunities in Open Generative AI](/content/2024/05/20/toward-responsible-innovation-evaluating-risks-and-opportunities-in-open-generative-ai/index.html)

- Mohammad Asjad

[TII Releases Falcon 2-11B: The First AI Model of the Falcon 2 Family Trained on 5.5T Tokens with a Vision Language Model](/content/2024/05/20/tii-releases-falcon-2-11b-the-first-ai-model-of-the-falcon-2-family-trained-on-5-5t-tokens-with-a-vision-language-model/index.html)

- Mohammad Asjad

[This AI Paper from Stanford University Evaluates the Performance of Multimodal Foundation Models Scaling from Few-Shot to Many-Shot-In-Context Learning ICL](/content/2024/05/19/this-ai-paper-from-stanford-university-evaluates-the-performance-of-multimodal-foundation-models-scaling-from-few-shot-to-many-shot-in-context-learning-icl/index.html)

- Mohammad Asjad

[Meta AI Introduces Chameleon: A New Family of Early-Fusion Token-based Foundation Models that Set a New Bar for Multimodal Machine Learning](/content/2024/05/18/meta-ai-introduces-chameleon-a-new-family-of-early-fusion-token-based-foundation-models-that-set-a-new-bar-for-multimodal-machine-learning/index.html)

- Mohammad Asjad

[SpeechVerse: A Multimodal AI Framework that Enables LLMs to Follow Natural Language Instructions for Performing Diverse Speech-Processing Tasks](/content/2024/05/17/speechverse-a-multimodal-ai-framework-that-enables-llms-to-follow-natural-language-instructions-for-performing-diverse-speech-processing-tasks/index.html)

- Mohammad Asjad

[Unveiling the Potential of Large Language Models: Enhancing Feedback Generation in Computing Education](/content/2024/05/16/unveiling-the-potential-of-large-language-models-enhancing-feedback-generation-in-computing-education/index.html)

- Mohammad Asjad

[CMU Researchers Propose MOMENT: A Family of Open-Source Machine Learning Foundation Models for General-Purpose Time Series Analysis](/content/2024/05/15/cmu-researchers-propose-moment-a-family-of-open-source-machine-learning-foundation-models-for-general-purpose-time-series-analysis/index.html)

- Mohammad Asjad

[Advancements in Knowledge Distillation and Multi-Teacher Learning: Introducing AM-RADIO Framework](/content/2024/05/15/advancements-in-knowledge-distillation-and-multi-teacher-learning-introducing-am-radio-framework/index.html)

- Mohammad Asjad

[RadOnc-GPT: Leveraging Meta Llama for a Pioneering Radiation Oncology Model](/content/2024/05/14/radonc-gpt-leveraging-meta-llama-for-a-pioneering-radiation-oncology-model/index.html)

- Mohammad Asjad

[Enhancing Anomaly Detection with Adaptive Noise: A Pseudo Anomaly Approach](/content/2024/05/13/enhancing-anomaly-detection-with-adaptive-noise-a-pseudo-anomaly-approach/index.html)

- Mohammad Asjad

[UC Berkeley Researchers Introduce Learnable Latent Codes as Bridges (LCB): A Novel AI Approach that Combines the Abstract Reasoning Capabilities of Large Language Models with Low-Level Action Policies](/content/2024/05/11/uc-berkeley-researchers-introduce-learnable-latent-codes-as-bridges-lcb-a-novel-ai-approach-that-combines-the-abstract-reasoning-capabilities-of-large-language-models-with-low-level-action-policies/index.html)

- Mohammad Asjad

[Towards Autonomous Software Development: The SWE-agent Revolution](/content/2024/05/10/towards-autonomous-software-development-the-swe-agent-revolution/index.html)

- Mohammad Asjad

[Exploring Sharpness-Aware Minimization (SAM): Insights into Label Noise Robustness and Generalization](/content/2024/05/09/exploring-sharpness-aware-minimization-sam-insights-into-label-noise-robustness-and-generalization/index.html)

- Mohammad Asjad

[Top AI-Powered Cartoonizer Tools](/content/2024/05/09/top-ai-powered-cartoonizer-tools/index.html)

- Mohammad Asjad

[TRAMBA: A Novel Hybrid Transformer and Mamba-based Architecture for Speech Super Resolution and Enhancement for Mobile and Wearable Platforms](/content/2024/05/08/tramba-a-novel-hybrid-transformer-and-mamba-based-architecture-for-speech-super-resolution-and-enhancement-for-mobile-and-wearable-platforms/index.html)

- Mohammad Asjad

[MaRDIFlow: Automating Metadata Abstraction for Enhanced Reproducibility in Computational Workflows](/content/2024/05/08/mardiflow-automating-metadata-abstraction-for-enhanced-reproducibility-in-computational-workflows/index.html)

- Mohammad Asjad

[Self-Play Preference Optimization (SPPO): An Innovative Machine Learning Approach to Finetuning Large Language Models (LLMs) from Human/AI Feedback](/content/2024/05/06/self-play-preference-optimization-sppo-an-innovative-machine-learning-approach-to-finetuning-large-language-models-llms-from-human-ai-feedback/index.html)

- Mohammad Asjad

[CMU Researchers Propose a Distributed Data Scoping Method: Revealing the Incompatibility between the Deep Learning Architecture and the Generic Transport PDEs](/content/2024/05/05/cmu-researchers-propose-a-distributed-data-scoping-method-revealing-the-incompatibility-between-the-deep-learning-architecture-and-the-generic-transport-pdes/index.html)

- Mohammad Asjad

[Deciphering Transformer Language Models: Advances in Interpretability Research](/content/2024/05/05/deciphering-transformer-language-models-advances-in-interpretability-research/index.html)

- Mohammad Asjad

[Researchers at Stanford Introduce SUQL: A Formal Query Language for Integrating Structured and Unstructured Data](/content/2024/05/04/researchers-at-stanford-introduce-suql-a-formal-query-language-for-integrating-structured-and-unstructured-data/index.html)

- Mohammad Asjad

[Evaluating LLM Trustworthiness: Insights from Harmoniticity Analysis Research from VISA Team](/content/2024/05/02/evaluating-llm-trustworthiness-insights-from-harmoniticity-analysis-research-from-visa-team/index.html)

- Mohammad Asjad

[Iterative Preference Optimization for Improving Reasoning Tasks in Language Models](/content/2024/05/02/iterative-preference-optimization-for-improving-reasoning-tasks-in-language-models/index.html)

- Mohammad Asjad

[Bridging the Binary Gap: Challenges in Training Neural Networks to Decode and Summarize Code](/content/2024/05/02/bridging-the-binary-gap-challenges-in-training-neural-networks-to-decode-and-summarize-code/index.html)

- Mohammad Asjad

[Meta AI Introduces CyberSecEval 2: A Novel Machine Learning Benchmark to Quantify LLM Security Risks and Capabilities](/content/2024/05/01/meta-ai-introduces-cyberseceval-2-a-novel-machine-learning-benchmark-to-quantify-llm-security-risks-and-capabilities/index.html)

- Mohammad Asjad

[Exploring Parameter-Efficient Fine-Tuning Strategies for Large Language Models](/content/2024/04/30/exploring-parameter-efficient-fine-tuning-strategies-for-large-language-models/index.html)

- Mohammad Asjad

[REBEL: A Reinforcement Learning RL Algorithm that Reduces the Problem of RL to Solving a Sequence of Relative Reward Regression Problems on Iteratively Collected Datasets](/content/2024/04/30/rebel-a-reinforcement-learning-rl-algorithm-that-reduces-the-problem-of-rl-to-solving-a-sequence-of-relative-reward-regression-problems-on-iteratively-collected-datasets/index.html)

- Mohammad Asjad

[Meet Electric Atlas: A New Era of Robotics by Boston Dynamics](/content/2024/04/30/meet-electric-atlas-a-new-era-of-robotics-by-boston-dynamics/index.html)

- Mohammad Asjad

[Researchers at UC San Diego Propose DrS: A Novel Machine Learning Approach for Learning Reusable Dense Rewards for Multi-Stage Tasks in a Data-Driven Manner](/content/2024/04/29/researchers-at-uc-san-diego-propose-drs-a-novel-machine-learning-approach-for-learning-reusable-dense-rewards-for-multi-stage-tasks-in-a-data-driven-manner/index.html)

- Mohammad Asjad

[From Lost to Found: INformation-INtensive (IN2) Training Revolutionizes Long-Context Language Understanding](/content/2024/04/28/from-lost-to-found-information-intensive-in2-training-revolutionizes-long-context-language-understanding/index.html)

- Mohammad Asjad

[Integrating Large Language Models with Graph Machine Learning: A Comprehensive Review](/content/2024/04/26/integrating-large-language-models-with-graph-machine-learning-a-comprehensive-review/index.html)

- Mohammad Asjad

[Enhancing AI Model’s Scalability and Performance: A Study on Multi-Head Mixture-of-Experts](/content/2024/04/25/enhancing-ai-models-scalability-and-performance-a-study-on-multi-head-mixture-of-experts/index.html)

- Mohammad Asjad

[Microsoft AI Releases Phi-3 Family of Models: A 3.8B Parameter Language Model Trained on 3.3T Tokens Locally on Your Phone](/content/2024/04/24/microsoft-ai-releases-phi-3-family-of-models-a-3-8b-parameter-language-model-trained-on-3-3t-tokens-locally-on-your-phone/index.html)

- Mohammad Asjad

[Interpretable Deep Learning for Biodiversity Monitoring: Introducing AudioProtoPNet](/content/2024/04/24/interpretable-deep-learning-for-biodiversity-monitoring-introducing-audioprotopnet/index.html)

- Mohammad Asjad

[Privacy-Preserving Training-as-a-Service (PTaaS): A Novel Service Computing Paradigm that Provides Privacy-Friendly and Customized Machine Learning Model Training for End Devices](/content/2024/04/23/privacy-preserving-training-as-a-service-ptaas-a-novel-service-computing-paradigm-that-provides-privacy-friendly-and-customized-machine-learning-model-training-for-end-devices/index.html)

- Mohammad Asjad

[NVIDIA AI Researchers Introduce ScaleFold: A Leap in High-Performance Computing for Protein Structure Prediction](/content/2024/04/22/nvidia-ai-researchers-introduce-scalefold-a-leap-in-high-performance-computing-for-protein-structure-prediction/index.html)

- Mohammad Asjad

[Unveiling Challenges in Language Model Performance: A Study of Saturation and Representation Degeneration](/content/2024/04/21/unveiling-challenges-in-language-model-performance-a-study-of-saturation-and-representation-degeneration/index.html)

- Mohammad Asjad

[Researchers at CMU Introduce TriForce: A Hierarchical Speculative Decoding AI System that is Scalable to Long Sequence Generation](/content/2024/04/20/researchers-at-cmu-introduce-triforce-a-hierarchical-speculative-decoding-ai-system-that-is-scalable-to-long-sequence-generation/index.html)

- Mohammad Asjad

[Researchers at Microsoft Introduces VASA-1: Transforming Realism in Talking Face Generation with Audio-Driven Innovation](/content/2024/04/19/researchers-at-microsoft-introduces-vasa-1-transforming-realism-in-talking-face-generation-with-audio-driven-innovation/index.html)

- Mohammad Asjad

[Unlocking the Recall Power of Large Language Models: Insights from Needle-in-a-Haystack Testing](/content/2024/04/19/unlocking-the-recall-power-of-large-language-models-insights-from-needle-in-a-haystack-testing/index.html)

- Mohammad Asjad

[Navigating the Landscape of CLIP: Investigating Data, Architecture, and Training Strategies](/content/2024/04/18/navigating-the-landscape-of-clip-investigating-data-architecture-and-training-strategies/index.html)

- Mohammad Asjad

[This AI Paper Introduces Pipeline Forward-Forward Algorithm (PFF): A Novel Machine Learning Approach to Training Distributed Neural Networks using Forward-Forward Algorithm](/content/2024/04/18/this-ai-paper-introduces-pipeline-forward-forward-algorithm-pff-a-novel-machine-learning-approach-to-training-distributed-neural-networks-using-forward-forward-algorithm/index.html)

- Mohammad Asjad

[Researchers at Oxford Presented Policy-Guided Diffusion: A Machine Learning Method for Controllable Generation of Synthetic Trajectories in Offline Reinforcement Learning RL](/content/2024/04/16/researchers-at-oxford-presented-policy-guided-diffusion-a-machine-learning-method-for-controllable-generation-of-synthetic-trajectories-in-offline-reinforcement-learning-rl/index.html)

- Mohammad Asjad

[Researchers at UC Berkeley Introduce GOEX: A Runtime for LLMs with an Intuitive Undo and Damage Confinement Abstractions, Enabling the Safer Deployment of LLM Agents in Practice](/content/2024/04/15/researchers-at-uc-berkeley-introduce-goex-a-runtime-for-llms-with-an-intuitive-undo-and-damage-confinement-abstractions-enabling-the-safer-deployment-of-llm-agents-in-practice/index.html)

- Mohammad Asjad

[LM-Guided CoT: A Novel Machine Learning Framework that Leverages a Lightweight (<1B) Language Model (LM) for guiding a black-box large (>10B) LM in Reasoning Tasks](/content/2024/04/15/lm-guided-cot-a-novel-machine-learning-framework-that-leverages-a-lightweight-1b-language-model-lm-for-guiding-a-black-box-large-10b-lm-in-reasoning-tasks/index.html)

- Mohammad Asjad

[The Future of Neural Network Training: Empirical Insights into μ-Transfer for Hyperparameter Scaling](/content/2024/04/13/the-future-of-neural-network-training-empirical-insights-into-%ce%bc-transfer-for-hyperparameter-scaling/index.html)

- Mohammad Asjad

[This AI Paper from China Introduces MiniCPM: Introducing Innovative Small Language Models Through Scalable Training Approaches](/content/2024/04/12/this-ai-paper-from-china-introduces-minicpm-introducing-innovative-small-language-models-through-scalable-training-approaches/index.html)

- Mohammad Asjad

[Meta AI Presents MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding](/content/2024/04/11/meta-ai-presents-ma-lmm-memory-augmented-large-multimodal-model-for-long-term-video-understanding/index.html)

- Mohammad Asjad

[This AI Paper Introduces ReasonEval: A New Machine Learning Method to Evaluate Mathematical Reasoning Beyond Accuracy](/content/2024/04/10/this-ai-paper-introduces-reasoneval-a-new-machine-learning-method-to-evaluate-mathematical-reasoning-beyond-accuracy/index.html)

- Mohammad Asjad

[This Machine Learning Paper Introduce PISSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models](/content/2024/04/10/this-machine-learning-paper-introduce-pissa-principal-singular-values-and-singular-vectors-adaptation-of-large-language-models/index.html)

- Mohammad Asjad

[Microsoft Researchers Propose Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models](/content/2024/04/09/microsoft-researchers-propose-visualization-of-thought-elicits-spatial-reasoning-in-large-language-models/index.html)

- Mohammad Asjad

[This Machine Learning Paper Introduces JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models](/content/2024/04/08/this-machine-learning-paper-introduces-jailbreakbench-an-open-robustness-benchmark-for-jailbreaking-large-language-models/index.html)

- Mohammad Asjad

[Linear Attention Sequence Parallel (LASP): An Efficient Machine Learning Method Tailored to Linear Attention-Based Language Models](/content/2024/04/07/linear-attention-sequence-parallel-lasp-an-efficient-machine-learning-method-tailored-to-linear-attention-based-language-models/index.html)

- Mohammad Asjad

[Researchers at Intel Labs Introduce LLaVA-Gemma: A Compact Vision-Language Model Leveraging the Gemma Large Language Model in Two Variants (Gemma-2B and Gemma-7B)](/content/2024/04/06/researchers-at-intel-labs-introduce-llava-gemma-a-compact-vision-language-model-leveraging-the-gemma-large-language-model-in-two-variants-gemma-2b-and-gemma-7b/index.html)

- Mohammad Asjad

[AutoTRIZ: An Artificial Ideation Tool that Leverages Large Language Models (LLMs) to Automate and Enhance the TRIZ (Theory of Inventive Problem Solving) Methodology](/content/2024/04/06/autotriz-an-artificial-ideation-tool-that-leverages-large-language-models-llms-to-automate-and-enhance-the-triz-theory-of-inventive-problem-solving-methodology/index.html)

- Mohammad Asjad

[UniLLMRec: An End-to-End LLM-Centered Recommendation Framework to Execute Multi-Stage Recommendation Tasks Through Chain-of-Recommendations](/content/2024/04/04/unillmrec-an-end-to-end-llm-centered-recommendation-framework-to-execute-multi-stage-recommendation-tasks-through-chain-of-recommendations/index.html)

- Mohammad Asjad

[Can Benign Data Undermine AI Safety? This Paper from Princeton University Explores the Paradox of Machine Learning Fine-Tuning](/content/2024/04/03/can-benign-data-undermine-ai-safety-this-paper-from-princeton-university-explores-the-paradox-of-machine-learning-fine-tuning/index.html)

- Mohammad Asjad

[Are We on the Right Way for Evaluating Large Vision-Language Models? This AI Paper from China Introduces MMStar: An Elite Vision-Dependent Multi-Modal Benchmark](/content/2024/04/03/are-we-on-the-right-way-for-evaluating-large-vision-language-models-this-ai-paper-from-china-introduces-mmstar-an-elite-vision-dependent-multi-modal-benchmark/index.html)

- Mohammad Asjad

[Evolution of RAGs: Naive RAG, Advanced RAG, and Modular RAG Architectures](/content/2024/04/01/evolution-of-rags-naive-rag-advanced-rag-and-modular-rag-architectures/index.html)

- Mohammad Asjad

[NVIDIA AI Research Proposes Language Instructed Temporal-Localization Assistant (LITA), which Enables Accurate Temporal Localization Using Video LLMs](/content/2024/03/31/nvidia-ai-research-proposes-language-instructed-temporal-localization-assistant-lita-which-enables-accurate-temporal-localization-using-video-llms/index.html)

- Mohammad Asjad

[Researchers from the University of Washington and Meta AI Present a Simple Context-Aware Decoding (CAD) Method to Encourage the Language Model to Attend to Its Context During Generation](/content/2024/03/30/researchers-from-the-university-of-washington-and-meta-ai-present-a-simple-context-aware-decoding-cad-method-to-encourage-the-language-model-to-attend-to-its-context-during-generation/index.html)

- Mohammad Asjad

[This Paper Reveals Insights from Reproducing OpenAI’s RLHF (Reinforcement Learning from Human Feedback) Work: Implementation and Scaling Explored](/content/2024/03/29/this-paper-reveals-insights-from-reproducing-openais-rlhf-reinforcement-learning-from-human-feedback-work-implementation-and-scaling-explored/index.html)

- Mohammad Asjad

[Do LLM Agents Have Regret? This Machine Learning Research from MIT and the University of Maryland Presents a Case Study on Online Learning and Games](/content/2024/03/28/do-llm-agents-have-regret-this-machine-learning-research-from-mit-and-the-university-of-maryland-presents-a-case-study-on-online-learning-and-games/index.html)

- Mohammad Asjad

[This AI Paper from Microsoft Present SiMBA: A Simplified Mamba-based Architecture for Vision and Multivariate Time Series](/content/2024/03/27/this-ai-paper-from-microsoft-present-simba-a-simplified-mamba-based-architecture-for-vision-and-multivariate-time-series/index.html)

- Mohammad Asjad

[Stability AI Introduces Stable Code: A General Purpose Base Code Language Model](/content/2024/03/27/stability-ai-introduces-stable-code-a-general-purpose-base-code-language-model/index.html)

- Mohammad Asjad

[The Idea of Compiler-Generated Feedback for Large Language Models](/content/2024/03/26/the-idea-of-compiler-generated-feedback-for-large-language-models/index.html)

- Mohammad Asjad

[LlamaFactory: A Unified Machine Learning Framework that Integrates a Suite of Cutting-Edge Efficient Training Methods, Allowing Users to Customize the Fine-Tuning of 100+ LLMs Flexibly](/content/2024/03/25/llamafactory-a-unified-machine-learning-framework-that-integrates-a-suite-of-cutting-edge-efficient-training-methods-allowing-users-to-customize-the-fine-tuning-of-100-llms-flexibly/index.html)

- Mohammad Asjad

[Sakana AI Introduces Evolutionary Model Merge: A New Machine Learning Approach Automating Foundation Model Development](/content/2024/03/24/sakana-ai-introduces-evolutionary-model-merge-a-new-machine-learning-approach-automating-foundation-model-development/index.html)

- Mohammad Asjad

[Researchers from Alibaba and the Renmin University of China Present mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding](/content/2024/03/23/researchers-from-alibaba-and-the-renmin-university-of-china-present-mplug-docowl-1-5-unified-structure-learning-for-ocr-free-document-understanding/index.html)

- Mohammad Asjad

[IBM’s Alignment Studio to Optimize AI Compliance for Contextual Regulations](/content/2024/03/22/ibms-alignment-studio-to-optimize-ai-compliance-for-contextual-regulations/index.html)

- Mohammad Asjad

[This AI Paper Introduces the Lightweight Mamba UNet (LightM-UNet) that Integrates Mamba and UNet in a Lightweight Framework for Medical Image Segmentation](/content/2024/03/18/this-ai-paper-introduces-the-lightweight-mamba-unet-lightm-unet-that-integrates-mamba-and-unet-in-a-lightweight-framework-for-medical-image-segmentation/index.html)

- Mohammad Asjad

[This Machine Learning Research Presents ScatterMoE: An Implementation of Sparse Mixture-of-Experts (SMoE) on GPUs](/content/2024/03/18/this-machine-learning-research-presents-scattermoe-an-implementation-of-sparse-mixture-of-experts-smoe-on-gpus/index.html)

- Mohammad Asjad

[Synth2: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings by Researchers from Google DeepMind](/content/2024/03/16/synth2-boosting-visual-language-models-with-synthetic-captions-and-image-embeddings-by-researchers-from-google-deepmind/index.html)

- Mohammad Asjad

[Google DeepMind Researchers Unveil Multistep Consistency Models: A Machine Learning Approach that Balances Speed and Quality in AI Sampling](/content/2024/03/14/google-deepmind-researchers-unveil-multistep-consistency-models-a-machine-learning-approach-that-balances-speed-and-quality-in-ai-sampling/index.html)

- Mohammad Asjad

[Training Value Functions via Classification for Scalable Deep Reinforcement Learning: Study by Google DeepMind Researchers and Others](/content/2024/03/11/training-value-functions-via-classification-for-scalable-deep-reinforcement-learning-study-by-google-deepmind-researchers-and-others/index.html)

- Mohammad Asjad

[This AI Paper from China Introduces ShortGPT: A Novel Artificial Intelligence Approach to Pruning Large Language Models (LLMs) based on Layer Redundancy](/content/2024/03/10/this-ai-paper-from-china-introduces-shortgpt-a-novel-artificial-intelligence-approach-to-pruning-large-language-models-llms-based-on-layer-redundancy/index.html)

- Mohammad Asjad

[Unlocking the Best Tokenization Strategies: How Greedy Inference and SaGe Lead the Way in NLP Models](/content/2024/03/09/unlocking-the-best-tokenization-strategies-how-greedy-inference-and-sage-lead-the-way-in-nlp-models/index.html)

- Mohammad Asjad

[Unlocking the ‘Wisdom of the Silicon Crowd’: How LLM Ensembles Are Redefining Forecasting Accuracy to Match Human Expertise](/content/2024/03/08/unlocking-the-wisdom-of-the-silicon-crowd-how-llm-ensembles-are-redefining-forecasting-accuracy-to-match-human-expertise/index.html)

- Mohammad Asjad

[Microsoft AI Researchers Developed a New Improved Framework ResLoRA for Low-Rank Adaptation (LoRA)](/content/2024/03/06/microsoft-ai-researchers-developed-a-new-improved-framework-reslora-for-low-rank-adaptation-lora/index.html)

- Mohammad Asjad

[Maximizing Efficiency in AI Training: A Deep Dive into Data Selection Practices and Future Directions](/content/2024/03/04/maximizing-efficiency-in-ai-training-a-deep-dive-into-data-selection-practices-and-future-directions/index.html)

- Mohammad Asjad

[Researchers from Tsinghua University and Microsoft AI Unveil a Breakthrough in Language Model Training: The Path to Optimal Learning Efficiency](/content/2024/03/04/researchers-from-tsinghua-university-and-microsoft-ai-unveil-a-breakthrough-in-language-model-training-the-path-to-optimal-learning-efficiency/index.html)

- Mohammad Asjad

[This AI Paper from the University of Michigan and Netflix Proposes CLoVe: A Machine Learning Framework to Improve the Compositionality of Pre-Trained Contrastive Vision-Language Models](/content/2024/03/03/this-ai-paper-from-the-university-of-michigan-and-netflix-proposes-clove-a-machine-learning-framework-to-improve-the-compositionality-of-pre-trained-contrastive-vision-language-models/index.html)

- Mohammad Asjad

[Researchers from Mohamed bin Zayed University of AI Developed ‘PALO’: A Polyglot Large Multimodal Model for 5B People](/content/2024/03/02/researchers-from-mohamed-bin-zayed-university-of-ai-developed-palo-a-polyglot-large-multimodal-model-for-5b-people/index.html)

- Mohammad Asjad

[Google AI Proposes USER-LLM: A Novel Artificial Intelligence Framework that Leverages User Embeddings to Contextualize LLMs](/content/2024/02/29/google-ai-proposes-user-llm-a-novel-artificial-intelligence-framework-that-leverages-user-embeddings-to-contextualize-llms/index.html)

- Mohammad Asjad

[Brown University Researchers Propose LexC-Gen: A New Artificial Intelligence Method that Generates Low-Resource-Language Classification Task Data at Scale](/content/2024/02/29/brown-university-researchers-propose-lexc-gen-a-new-artificial-intelligence-method-that-generates-low-resource-language-classification-task-data-at-scale/index.html)

- Mohammad Asjad

[Neural Network Diffusion: Generating High-Performing Neural Network Parameters](/content/2024/02/28/neural-network-diffusion-generating-high-performing-neural-network-parameters/index.html)

- Mohammad Asjad

[Microsoft Present AI Controller Interface: Generative AI with a Lightweight, LLM-Integrated Virtual Machine (VM)](/content/2024/02/27/microsoft-present-ai-controller-interface-generative-ai-with-a-lightweight-llm-integrated-virtual-machine-vm/index.html)

- Mohammad Asjad

[Can We Drastically Reduce AI Training Costs? This AI Paper from MIT, Princeton, and Together AI Unveils How BitDelta Achieves Groundbreaking Efficiency in Machine Learning](/content/2024/02/26/can-we-drastically-reduce-ai-training-costs-this-ai-paper-from-mit-princeton-and-together-ai-unveils-how-bitdelta-achieves-groundbreaking-efficiency-in-machine-learning/index.html)

- Mohammad Asjad

[Can Machine Learning Teach Robots to Understand Us Better? This Microsoft Research Introduces Language Feedback Models for Advanced Imitation Learning](/content/2024/02/25/can-machine-learning-teach-robots-to-understand-us-better-this-microsoft-research-introduces-language-feedback-models-for-advanced-imitation-learning/index.html)

- Mohammad Asjad

[This Machine Learning Research from Yale and Google AI Introduce SubGen: An Efficient Key-Value Cache Compression Algorithm via Stream Clustering](/content/2024/02/23/this-machine-learning-research-from-yale-and-google-ai-introduce-subgen-an-efficient-key-value-cache-compression-algorithm-via-stream-clustering/index.html)

- Mohammad Asjad

[Apple Researchers Introduce Keyframer: An LLM-Powered Animation Prototyping Tool that can Generate Animations from Static Images (SVGs)](/content/2024/02/22/apple-researchers-introduce-keyframer-an-llm-powered-animation-prototyping-tool-that-can-generate-animations-from-static-images-svgs/index.html)

- Mohammad Asjad

[Meet BiLLM: A Novel Post-Training Binary Quantization Method Specifically Tailored for Compressing Pre-Trained LLMs](/content/2024/02/20/meet-billm-a-novel-post-training-binary-quantization-method-specifically-tailored-for-compressing-pre-trained-llms/index.html)

- Mohammad Asjad

[This AI Paper Proposes an Interactive Agent Foundation Model that Uses a Novel Multi-Task Agent Training Paradigm for Training AI Agents Across a Wide Range of Domains, Datasets, and Tasks](/content/2024/02/17/this-ai-paper-proposes-an-interactive-agent-foundation-model-that-uses-a-novel-multi-task-agent-training-paradigm-for-training-ai-agents-across-a-wide-range-of-domains-datasets-and-tasks/index.html)

- Mohammad Asjad

[Meet EscherNet: A Multi-View Conditioned Diffusion Model for View Synthesis](/content/2024/02/14/meet-eschernet-a-multi-view-conditioned-diffusion-model-for-view-synthesis/index.html)

- Mohammad Asjad

[Can Large Language Models be Trusted for Evaluation? Meet SCALEEVAL: An Agent-Debate-Assisted Meta-Evaluation Framework that Leverages the Capabilities of Multiple Communicative LLM Agents](/content/2024/02/11/can-large-language-models-be-trusted-for-evaluation-meet-scaleeval-an-agent-debate-assisted-meta-evaluation-framework-that-leverages-the-capabilities-of-multiple-communicative-llm-agents/index.html)

- Mohammad Asjad

[Stanford Researchers Introduce RAPTOR: A Novel Tree-based Retrieval System that Augments the Parametric Knowledge of LLMs with Contextual Information](/content/2024/02/08/stanford-researchers-introduce-raptor-a-novel-tree-based-retrieval-system-that-augments-the-parametric-knowledge-of-llms-with-contextual-information/index.html)

- Mohammad Asjad

[Researchers from McGill University Present the Pythia 70M Model for Distilling Transformers into Long Convolution Models](/content/2024/02/07/researchers-from-mcgill-university-present-the-pythia-70m-model-for-distilling-transformers-into-long-convolution-models/index.html)

- Mohammad Asjad

[Alibaba Researchers Introduce Mobile-Agent: An Autonomous Multi-Modal Mobile Device Agent](/content/2024/02/04/alibaba-researchers-introduce-mobile-agent-an-autonomous-multi-modal-mobile-device-agent/index.html)

- Mohammad Asjad

[Researchers from the Chinese University of Hong Kong and Tencent AI Lab Propose a Multimodal Pathway to Improve Transformers with Irrelevant Data from Other Modalities](/content/2024/02/01/researchers-from-the-chinese-university-of-hong-kong-and-tencent-ai-lab-propose-a-multimodal-pathway-to-improve-transformers-with-irrelevant-data-from-other-modalities/index.html)

- Mohammad Asjad

[This AI Paper from China Introduces DREditor: A Time-Efficient AI Approach for Building a Domain-Specific Dense Retrieval Model](/content/2024/01/30/this-ai-paper-from-china-introduces-dreditor-a-time-efficient-ai-approach-for-building-a-domain-specific-dense-retrieval-model/index.html)

- Mohammad Asjad

[This AI Paper from ETH Zurich, Google, and Max Plank Proposes an Effective AI Strategy to Boost the Performance of Reward Models for RLHF (Reinforcement Learning from Human Feedback)](/content/2024/01/27/this-ai-paper-from-eth-zurich-google-and-max-plank-proposes-an-effective-ai-strategy-to-boost-the-performance-of-reward-models-for-rlhf-reinforcement-learning-from-human-feedback/index.html)

- Mohammad Asjad

[Google AI Presents Lumiere: A Space-Time Diffusion Model for Video Generation](/content/2024/01/26/google-ai-presents-lumiere-a-space-time-diffusion-model-for-video-generation/index.html)

- Mohammad Asjad

[Researchers from ByteDance and Sun Yat-Sen University Introduce DiffusionGPT: LLM-Driven Text-to-Image Generation System](/content/2024/01/24/researchers-from-bytedance-and-sun-yat-sen-university-introduce-diffusiongpt-llm-driven-text-to-image-generation-system/index.html)

- Mohammad Asjad

[This AI Paper from Meta and NYU Introduces Self-Rewarding Language Models that are Capable of Self-Alignment via Judging and Training on their Own Generations](/content/2024/01/22/this-ai-paper-from-meta-and-nyu-introduces-self-rewarding-language-models-that-are-capable-of-self-alignment-via-judging-and-training-on-their-own-generations/index.html)

- Mohammad Asjad

[Researchers from the University of Washington and Allen Institute for AI Present Proxy-Tuning: An Efficient Alternative to Finetuning Large Language Models](/content/2024/01/21/researchers-from-the-university-of-washington-and-allen-institute-for-ai-present-proxy-tuning-an-efficient-alternative-to-finetuning-large-language-models/index.html)

- Mohammad Asjad

[This AI Paper Introduces XAI-AGE: A Groundbreaking Deep Neural Network for Biological Age Prediction and Insight into Epigenetic Mechanisms](/content/2024/01/19/this-ai-paper-introduces-xai-age-a-groundbreaking-deep-neural-network-for-biological-age-prediction-and-insight-into-epigenetic-mechanisms/index.html)

- Mohammad Asjad

[Stanford Researchers Introduce Clover: Closed-Loop Verifiable Code Generation that Checks Consistencies Among Code, Doc Strings and Annotations and Enforces Correctness in AI-Generated Code](/content/2024/01/16/stanford-researchers-introduce-clover-closed-loop-verifiable-code-generation-that-checks-consistencies-among-code-doc-strings-and-annotations-and-enforces-correctness-in-ai-generated/index.html)

- Mohammad Asjad

[This AI Paper from China Unveils ‘Activation Beacon’: A Groundbreaking AI Technique to Expand Context Understanding in Large Language Models](/content/2024/01/15/this-ai-paper-from-china-unveils-activation-beacon-a-groundbreaking-ai-technique-to-expand-context-understanding-in-large-language-models/index.html)

- Mohammad Asjad

[AWS Researchers Propose Panda: A New Machine Learning Framework to Provide Context Grounding to Pre-Trained LLMs](/content/2024/01/14/aws-researchers-propose-panda-a-new-machine-learning-framework-to-provide-context-grounding-to-pre-trained-llms/index.html)

- Mohammad Asjad

[NTU and Meta Researchers Introduce URHand: A Universal Relightable Hand AI Model that Generalizes Across Viewpoints, Poses, Illuminations, and Identities](/content/2024/01/12/ntu-and-meta-researchers-introduce-urhand-a-universal-relightable-hand-ai-model-that-generalizes-across-viewpoints-poses-illuminations-and-identities/index.html)

#### [RELATED ARTICLES](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/\#/index.html) [MORE FROM AUTHOR](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/\#/index.html)

### [How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)

### [Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order](/content/2026/06/13/anthropic-disables-claude-fable-5-and-mythos-5-after-us-government-order/ "Anthropic Disables Claude Fable 5 and Mythos 5 After US Government Order"/index.html)

### [Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)

### [A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)

### [Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard](/content/2026/06/12/google-releases-gemini-sql2-gemini-3-1-pro-text-to-sql-scores-80-04-on-bird-single-model-leaderboard/ "Google Releases Gemini-SQL2: Gemini 3.1 Pro Text-to-SQL Scores 80.04% on BIRD Single-Model Leaderboard"/index.html)

### [Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)

[prev-page](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/#/index.html)[next-page](/content/2024/07/25/lean-github-a-large-scale-dataset-for-advancing-automated-theorem-proving/#/index.html)

### [How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access,...](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/ "How to Build a QwenPaw Agent Workspace with Custom Skills, Model Providers, Console Access, and Streaming API Testing"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 13, 2026[0](/content/2026/06/13/how-to-build-a-qwenpaw-agent-workspace-with-custom-skills-model-providers-console-access-and-streaming-api-testing/#respond/index.html)

In this tutorial, we implement a QwenPaw workflow that provides a practical environment for building and testing an agent-powered assistant. We install and initialize...

[Asif Razzaq](/content/author/6flvq/index.html)-June 13, 2026[0](/content/2026/06/13/anthropic-disables-claude-fable-5-and-mythos-5-after-us-government-order/#respond/index.html)

shutdown followed a US government export control directive citing national security authorities. All other Anthropic models, including Opus 4.8, remain available.

### [Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench...](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/#respond/index.html)

Moonshot AI has open-sourced Kimi K2.7-Code under a Modified MIT license. It is a coding-focused, agentic model built on Kimi K2.6, with a 256K context window and roughly 30% lower reasoning-token usage. Moonshot reports gains over K2.6 on six benchmarks, including +21.8% on Kimi Code Bench v2. The model is available via the Kimi API and Kimi Code.

### [A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph,...](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/ "A Coding Implementation on Spatial Graph Neural Networks for Urban Function Inference Using city2graph, OSMnx, and PyTorch Geometric"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 12, 2026[0](/content/2026/06/12/a-coding-implementation-on-spatial-graph-neural-networks-for-urban-function-inference-using-city2graph-osmnx-and-pytorch-geometric/#respond/index.html)

We build an end-to-end spatial graph learning pipeline using city2graph. We collect urban POI and street network data from OpenStreetMap, with a synthetic fallback for reliability. We engineer spatial features, construct several proximity graph families, and compare how each represents the same urban environment. We then build heterogeneous and homogeneous graphs, convert them to PyTorch Geometric, and train a GraphSAGE model to predict POI categories from spatial structure.

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/google-releases-gemini-sql2-gemini-3-1-pro-text-to-sql-scores-80-04-on-bird-single-model-leaderboard/#respond/index.html)

We look at Gemini-SQL2, the text-to-SQL capability Google Research announced on June 12, 2026. Powered by Gemini 3.1 Pro, it posted 80.04% execution accuracy on the BIRD single-model leaderboard. We explain what the score measures, how the leaderboard stacks up, and what Google has not yet disclosed. We also cover use cases and a schema-grounded implementation pattern.

### [Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6...](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/ "Moonshot AI Launches Kimi Work, a Local Desktop Agent Reportedly Running on Kimi K2.6 With a 300-Sub-Agent Agent Swarm"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/moonshot-ai-launches-kimi-work-a-local-desktop-agent-reportedly-running-on-kimi-k2-6-with-a-300-sub-agent-agent-swarm/#respond/index.html)

Moonshot AI's Kimi Work is a local desktop agent for macOS and Windows. It runs a 300-sub-agent swarm, drives your logged-in browser via WebBridge, and schedules background jobs.

### [Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order...](/content/2026/06/12/zyphra-release-zamba2-vl-hybrid-mamba2-transformer-vision-language-models-that-cut-time-to-first-token-by-about-an-order-of-magnitude/ "Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude"/index.html)

[Asif Razzaq](/content/author/6flvq/index.html)-June 12, 2026[0](/content/2026/06/12/zyphra-release-zamba2-vl-hybrid-mamba2-transformer-vision-language-models-that-cut-time-to-first-token-by-about-an-order-of-magnitude/#respond/index.html)

Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about an order of magnitude.

### [A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical...](/content/2026/06/12/a-coding-implementation-on-monai-for-end-to-end-3d-spleen-segmentation-using-unet-on-medical-ct-volumes/ "A Coding Implementation on MONAI for End-to-End 3D Spleen Segmentation Using UNet on Medical CT Volumes"/index.html)

[Sana Hassan](/content/author/sana-hassan/index.html)-June 12, 2026[0](/content/2026/06/12/a-coding-implementation-on-monai-for-end-to-end-3d-spleen-segmentation-using-unet-on-medical-ct-volumes/#respond/index.html)

In this tutorial, we build an end-to-end 3D medical image segmentation pipeline using MONAI to segment the spleen on the Medical Segmentation Decathlon Task09...

### [Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For...](/content/2026/06/11/perplexity-moves-deep-research-into-computer-routing-research-subtasks-across-20-frontier-models-for-reports-decks-and-dashboards/ "Perplexity Moves Deep Research Into Computer, Routing Research Subtasks Across 20+ Frontier Models For Reports, Decks, And Dashboards"/index.html)

[Michal Sutter](/content/author/michal-sutter/index.html)-June 11, 2026[0](/content/2026/06/11/perplexity-moves-deep-research-into-computer-routing-research-subtasks-across-20-frontier-models-for-reports-decks-and-dashboards/#respond/index.html)

Deep Research now lives inside Perplexity Computer, breaking hard questions into subtasks and routing across 20+ frontier models.

### [xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and...](/content/2026/06/11/xai-ships-grok-build-plugin-marketplace-with-mongodb-vercel-sentry-chrome-devtools-cloudflare-and-superpowers-plugins-at-launch/ "xAI Ships Grok Build Plugin Marketplace With MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and Superpowers Plugins at Launch"/index.html)

[Michal Sutter](/content/author/michal-sutter/index.html)-June 11, 2026[0](/content/2026/06/11/xai-ships-grok-build-plugin-marketplace-with-mongodb-vercel-sentry-chrome-devtools-cloudflare-and-superpowers-plugins-at-launch/#respond/index.html)

Grok Build's in-terminal marketplace bundles skills, agents, hooks, and MCP servers, with commit-SHA verification on every remote plugin.

- [miniCON Event 2025](https://pxl.to/hki7r39)
- [Download](/content/download/index.html)
  - [AI Magazine/Report](/content/ai-magazine/index.html)
- [Privacy & TC](/content/privacy-policy/index.html)
- [Cookie Policy](/content/cookie-policy/index.html)
- [Newsletter](https://www.aidevsignals.com/)
- [Partnership and Promotion](https://forms.gle/mjneG2kKPjDu6Hv8A)

© Copyright Reserved @2025 Marktechpost AI Media Inc
