- Популярные видео
- Авто
- Видео-блоги
- ДТП, аварии
- Для маленьких
- Еда, напитки
- Животные
- Закон и право
- Знаменитости
- Игры
- Искусство
- Комедии
- Красота, мода
- Кулинария, рецепты
- Люди
- Мото
- Музыка
- Мультфильмы
- Наука, технологии
- Новости
- Образование
- Политика
- Праздники
- Приколы
- Природа
- Происшествия
- Путешествия
- Развлечения
- Ржач
- Семья
- Сериалы
- Спорт
- Стиль жизни
- ТВ передачи
- Танцы
- Технологии
- Товары
- Ужасы
- Фильмы
- Шоу-бизнес
- Юмор
Erfan Shayegani - Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
00:00 Intro and Paper Overview
00:46 What Are Computer Use Agents
02:50 Safety Focus and Blind Goal Directedness
07:37 Pattern One Context Failures
13:45 Pattern Two Risky Assumptions
17:47 Pattern Three Infeasible Goals
20:02 BlindAct Benchmark Setup
23:45 Evaluation With LLM Judges
25:20 Results BGD vs Completion
31:41 Prompting Mitigations Limits
35:03 Qualitative Failure Modes
39:51 Q&A and Closing
In this talk Erfan will talk about adversarial attacks on multi-modal language models.
Computer-Use Agents (CUAs) are an increasingly deployed class of agents that take actions on GUIs to accomplish user goals. In this paper, we show that CUAs consistently exhibit Blind Goal-Directedness (BGD): a bias to pursue goals regardless of feasibility, safety, reliability, or context. We characterize three prevalent patterns of BGD: (i) lack of contextual reasoning, (ii) assumptions and decisions under ambiguity, and (iii) contradictory or infeasible goals. We develop BLIND-ACT, a benchmark of 90 tasks capturing these three patterns. Built on OSWorld, BLIND-ACT provides realistic environments and employs LLM-based judges to evaluate agent behavior, achieving 93.75% agreement with human annotations. We use BLIND-ACT to evaluate nine frontier models, including Claude Sonnet and Opus 4, Computer-Use-Preview, and GPT-5, observing high average BGD rates (80.8%) across them. We show that BGD exposes subtle risks that arise even when inputs are not directly harmful. While prompting-based interventions lower BGD levels, substantial risk persists, highlighting the need for stronger training- or inference-time interventions. Qualitative analysis reveals observed failure modes: execution-first bias (focusing on how to act over whether to act), thought–action disconnect (execution diverging from reasoning), and request-primacy (justifying actions due to user request). Identifying BGD and introducing BLIND-ACT establishes a foundation for future research on studying and mitigating this fundamental risk and ensuring safe CUA deployment. The paper has been accepted to ICLR 2026 with Microsoft Research.
Erfan is a 4th-year PhD student at the University of California, Riverside. His research focuses on the intersection of Generative AI and trustworthiness, particularly on Multimodal Language Models (LLMs/MLLMs) and AI Agents such as Computer-Use Agents (CUAs) with an emphasis on Alignment, Robustness, Safety, Ethics, Fairness, Bias, and Security/Privacy. Erfan is driven by an adversarial mindset, probing models for alignment gaps to build more robust systems.
As a two-time Microsoft Research intern, he has pivoted between red-teaming "computer-use agents" for ICLR 2026 and developing patented methods to steer empathy and user satisfaction in LLMs.
This session is brought to you by the Cohere Labs Open Science Community - a space where
ML researchers, engineers, linguists, social scientists, and lifelong learners connect and collaborate with each other. We'd like to extend a special thank you to Manuel Villanueva and Damani
Leads of ourPrivacy, Security and Policy group for their dedication in organizing this event.
If you’re interested in sharing your work, we welcome you to join us! Simply fill out the form at https://forms.gle/ALND9i6KouEEpCnz6 to express your interest in becoming a speaker.
Join the Cohere Labs Open Science Community to see a full list of upcoming events (https://tinyurl.com/CohereLabsCommunityApp).
Видео Erfan Shayegani - Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness канала Cohere
00:46 What Are Computer Use Agents
02:50 Safety Focus and Blind Goal Directedness
07:37 Pattern One Context Failures
13:45 Pattern Two Risky Assumptions
17:47 Pattern Three Infeasible Goals
20:02 BlindAct Benchmark Setup
23:45 Evaluation With LLM Judges
25:20 Results BGD vs Completion
31:41 Prompting Mitigations Limits
35:03 Qualitative Failure Modes
39:51 Q&A and Closing
In this talk Erfan will talk about adversarial attacks on multi-modal language models.
Computer-Use Agents (CUAs) are an increasingly deployed class of agents that take actions on GUIs to accomplish user goals. In this paper, we show that CUAs consistently exhibit Blind Goal-Directedness (BGD): a bias to pursue goals regardless of feasibility, safety, reliability, or context. We characterize three prevalent patterns of BGD: (i) lack of contextual reasoning, (ii) assumptions and decisions under ambiguity, and (iii) contradictory or infeasible goals. We develop BLIND-ACT, a benchmark of 90 tasks capturing these three patterns. Built on OSWorld, BLIND-ACT provides realistic environments and employs LLM-based judges to evaluate agent behavior, achieving 93.75% agreement with human annotations. We use BLIND-ACT to evaluate nine frontier models, including Claude Sonnet and Opus 4, Computer-Use-Preview, and GPT-5, observing high average BGD rates (80.8%) across them. We show that BGD exposes subtle risks that arise even when inputs are not directly harmful. While prompting-based interventions lower BGD levels, substantial risk persists, highlighting the need for stronger training- or inference-time interventions. Qualitative analysis reveals observed failure modes: execution-first bias (focusing on how to act over whether to act), thought–action disconnect (execution diverging from reasoning), and request-primacy (justifying actions due to user request). Identifying BGD and introducing BLIND-ACT establishes a foundation for future research on studying and mitigating this fundamental risk and ensuring safe CUA deployment. The paper has been accepted to ICLR 2026 with Microsoft Research.
Erfan is a 4th-year PhD student at the University of California, Riverside. His research focuses on the intersection of Generative AI and trustworthiness, particularly on Multimodal Language Models (LLMs/MLLMs) and AI Agents such as Computer-Use Agents (CUAs) with an emphasis on Alignment, Robustness, Safety, Ethics, Fairness, Bias, and Security/Privacy. Erfan is driven by an adversarial mindset, probing models for alignment gaps to build more robust systems.
As a two-time Microsoft Research intern, he has pivoted between red-teaming "computer-use agents" for ICLR 2026 and developing patented methods to steer empathy and user satisfaction in LLMs.
This session is brought to you by the Cohere Labs Open Science Community - a space where
ML researchers, engineers, linguists, social scientists, and lifelong learners connect and collaborate with each other. We'd like to extend a special thank you to Manuel Villanueva and Damani
Leads of ourPrivacy, Security and Policy group for their dedication in organizing this event.
If you’re interested in sharing your work, we welcome you to join us! Simply fill out the form at https://forms.gle/ALND9i6KouEEpCnz6 to express your interest in becoming a speaker.
Join the Cohere Labs Open Science Community to see a full list of upcoming events (https://tinyurl.com/CohereLabsCommunityApp).
Видео Erfan Shayegani - Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness канала Cohere
Комментарии отсутствуют
Информация о видео
18 мая 2026 г. 23:01:37
00:50:24
Другие видео канала
