`
- inside the `qbsearch-input`. I will try clicking [653] first.
- page_context_url: https://github.com/
- page_context_title: "GitHub \xB7 Change is constant. GitHub keeps you ahead. \xB7\
- \ GitHub"
-- description: Press Enter key
- output: null
- cssSelector: null
- xpath: null
- elementTag: null
- elementHash: null
- selectorStrategies: null
- target_text: element
- container_hint: null
- position_hint: null
- interaction_type: null
- type: key_press
- key: Enter
-- description: Scroll down 1.0 pages
- output: null
- type: scroll
- scrollX: 0
- scrollY: 800
- agent_reasoning: The previous step of inputting "browser-use" and pressing Enter
- was successful, and I am now on the search results page. The immediate goal is
- to find and click the "browser-use" repository link. Since the search was executed,
- the current page should display the results. I need to scroll down to see the
- search results or use `find_text` to locate the repository link. Since the current
- view is still showing the search modal/overlay, I will scroll down to see the
- results page content.
- page_context_url: https://github.com/
-- description: Scroll down 1.0 pages
- output: null
- type: scroll
- scrollX: 0
- scrollY: 800
- agent_reasoning: The previous step of scrolling down was performed because the search
- results were expected to be below the current view, which was still showing the
- search modal/overlay after pressing Enter. However, the current browser state
- still shows the search overlay elements (like the query builder and suggestions)
- at the top, and the main content of the search results page is still below. I
- need to scroll further down to see the search results, or the search result page
- might not have loaded correctly after the Enter key, or the scroll didn't move
- the viewport enough relative to the search results. I will try scrolling down
- again to ensure the search results are visible. The page info shows 11.1 pages
- below, so scrolling is necessary.
- page_context_url: https://github.com/
-- description: Scroll down 5.0 pages
- output: null
- type: scroll
- scrollX: 0
- scrollY: 4000
- agent_reasoning: The previous two steps involved scrolling down, but the browser
- state still shows the search overlay prominently, and the search results are not
- yet visible. I need to scroll down significantly more to get past the initial
- landing page content and see the search results page content. The page info suggests
- 11.1 pages below, so I will scroll down by a larger amount (e.g., 5 pages) to
- ensure I reach the search results.
- page_context_url: https://github.com/
-- description: Scroll down 10.0 pages
- output: null
- type: scroll
- scrollX: 0
- scrollY: 8000
- agent_reasoning: The previous steps involved navigating to GitHub, searching for
- "browser-use", and repeatedly scrolling down. Despite three attempts at scrolling,
- the browser state still shows the search overlay and the initial landing page
- content, not the actual search results page content. This suggests that either
- the search was not executed correctly, or the search results are very far down,
- or the page structure is confusing and the search results are hidden. Since I
- pressed Enter after inputting the text, the search should have executed. I will
- try scrolling down one more time, but this time I will use `scroll(down=True,
- pages=10)` to ensure I reach the bottom of the page and see the search results.
- If that fails, I will suspect the search execution failed and try to find a "Search"
- button to click instead of relying on the Enter key, or re-evaluate the search
- input. Given the page info shows 11.1 pages below, scrolling to the bottom is
- the next logical step to confirm the search results are loaded.
- page_context_url: https://github.com/
-- description: Click on 'browser-use, Search all of GitHub' (News/Press Releases)
- output: null
- cssSelector: ''
- xpath: html/body/div[1]/div[3]/header/div/div[2]/div/div/qbsearch-input/div[1]/div/modal-dialog/div/div/div/form/query-builder/div[1]/div[2]/ul/li/ul/li
- elementTag: li
- elementHash: abaa6b85ca
- selectorStrategies: null
- target_text: browser-use, Search all of GitHub
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: 'Despite multiple attempts to scroll down, the search results page
- content is still not visible, and the search overlay remains prominent. This strongly
- suggests that either the search was not successfully executed (i.e., pressing
- Enter in the shadow DOM input didn''t trigger the navigation), or the search results
- are not on the current page but require a click on a specific search result link
- within the overlay to navigate to the results page.
-
-
- Looking at the current interactive elements, I see a suggestion: `[10021]
`. This list item represents the search result/suggestion
- "browser-use, Search all of GitHub". I should click this element to execute the
- search and navigate to the results page.'
- page_context_url: https://github.com/
- page_context_title: "GitHub \xB7 Change is constant. GitHub keeps you ahead. \xB7\
- \ GitHub"
-- description: "Click on Browser Use (on Repository search results \xB7 GitHub)"
- output: null
- cssSelector: ''
- xpath: html/body/div[1]/div[4]/main/react-app/div/div/div[1]/div/div/div[2]/div/div/div[1]/div[4]/div/div/div[1]/div/div[1]/h3/div/div[2]/a
- elementTag: a
- elementHash: 7a2302537b
- selectorStrategies: null
- target_text: Browser Use
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: The previous step of clicking the search suggestion successfully
- navigated to the repository search results page for "browser-use". I can confirm
- this by the current URL and the page title. The next immediate goal is to click
- on the correct repository link, which is "browser-use/browser-use" based on the
- user request and the search results list. I see element [14628] which is an `
`
- tag with the text "browser-use/ browser-use". This is the target repository. After
- clicking, I will be on the repository page and can extract the star count.
- page_context_url: https://github.com/
- page_context_title: "Repository search results \xB7 GitHub"
-input_schema: []
diff --git a/workflows/storage/workflows/bfaf87d9-f5d9-4245-b210-5bfdf302a9b2.workflow.yaml b/workflows/storage/workflows/bfaf87d9-f5d9-4245-b210-5bfdf302a9b2.workflow.yaml
deleted file mode 100644
index eee23dc3..00000000
--- a/workflows/storage/workflows/bfaf87d9-f5d9-4245-b210-5bfdf302a9b2.workflow.yaml
+++ /dev/null
@@ -1,135 +0,0 @@
-workflow_analysis: null
-name: '
-
- Go to GitHub, search for the browser-use repository, click on it,
-
- and extract the star count.
-
- '
-description: "Workflow for: \nGo to GitHub, search for the browser-use repository,\
- \ click on it,\nand extract the star count.\n"
-version: 1.0.0
-steps:
-- description: Navigate to https://github.com
- output: null
- type: navigation
- url: https://github.com
- agent_reasoning: The first step is to navigate to GitHub. Since the current page
- is empty, I will use the `navigate` action. The ultimate goal is to find the "browser-use"
- repository and extract its star count.
-- description: "Click on 'Search or jump to\u2026' (Filings section)"
- output: null
- cssSelector: ''
- xpath: html/body/div[1]/div[3]/header/div/div[2]/div/div/qbsearch-input/div[1]/button
- elementTag: button
- elementHash: e53e5f0519
- selectorStrategies: null
- target_text: "Search or jump to\u2026"
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: "The previous step of waiting for 3 seconds was successful, and\
- \ the page content is now visible. The immediate goal is to find the search bar\
- \ and input \"browser-use\". I see an interactive element `[520] ` which looks like the search\
- \ button/input. It is wrapped in a `qbsearch-input` element. I will click this\
- \ button to activate the search input, and then input the search query."
- page_context_url: https://github.com/
- page_context_title: "GitHub \xB7 Change is constant. GitHub keeps you ahead. \xB7\
- \ GitHub"
-- description: Press browser-use{Enter} key
- output: null
- cssSelector: null
- xpath: null
- elementTag: null
- elementHash: null
- selectorStrategies: null
- target_text: element
- container_hint: null
- position_hint: null
- interaction_type: null
- type: key_press
- key: browser-use{Enter}
-- description: "Click on browser-use{Enter}, Search all of GitHub (on GitHub \xB7\
- \ Change is constant. GitHub keeps you ahead. \xB7 GitHub)"
- output: null
- cssSelector: ''
- xpath: html/body/div[1]/div[3]/header/div/div[2]/div/div/qbsearch-input/div[1]/div/modal-dialog/div/div/div/form/query-builder/div[1]/div[2]/ul/li/ul/li
- elementTag: li
- elementHash: abaa6b85ca
- selectorStrategies: null
- target_text: browser-use{Enter}, Search all of GitHub
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: 'The previous step used `send_keys("browser-use{Enter}")` as a
- last resort to input the search query after repeated failures with `evaluate()`
- on the shadow DOM input. The current browser state shows the text "browser-use{Enter}"
- has been entered into the search box, but the search has not yet been executed
- (the page hasn''t changed to the search results). The input field is still visible.
- I need to trigger the search. Since the `send_keys` included `{Enter}`, the search
- should have been submitted, but it seems GitHub''s search mechanism requires a
- click on a suggestion or a specific search button if the input is inside a complex
- component.
-
-
- Looking at the current state, there is a suggestion `[7052]` which says `browser-use{Enter},
- Search all of GitHub`. This suggests the text was entered but the search wasn''t
- executed. I will try to click the suggestion `[7052]` to execute the search.'
- page_context_url: https://github.com/
- page_context_title: "GitHub \xB7 Change is constant. GitHub keeps you ahead. \xB7\
- \ GitHub"
-- description: "Click on browser-use{Enter}, Search all of GitHub (on GitHub \xB7\
- \ Change is constant. GitHub keeps you ahead. \xB7 GitHub)"
- output: null
- cssSelector: ''
- xpath: html/body/div[1]/div[3]/header/div/div[2]/div/div/qbsearch-input/div[1]/div/modal-dialog/div/div/div/form/query-builder/div[1]/div[2]/ul/li/ul/li
- elementTag: li
- elementHash: abaa6b85ca
- selectorStrategies: null
- target_text: browser-use{Enter}, Search all of GitHub
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: 'The previous step used `send_keys("browser-use{Enter}")` as a
- last resort to input the search query after repeated failures with `evaluate()`
- on the shadow DOM input. The current browser state shows the text "browser-use{Enter}"
- has been entered into the search box, but the search has not yet been executed
- (the page hasn''t changed to the search results). The input field is still visible.
- I need to trigger the search. Since the `send_keys` included `{Enter}`, the search
- should have been submitted, but it seems GitHub''s search mechanism requires a
- click on a suggestion or a specific search button if the input is inside a complex
- component.
-
-
- Looking at the current state, there is a suggestion `[7052]` which says `browser-use{Enter},
- Search all of GitHub`. This suggests the text was entered but the search wasn''t
- executed. I will try to click the suggestion `[7052]` to execute the search.'
- page_context_url: https://github.com/
- page_context_title: "GitHub \xB7 Change is constant. GitHub keeps you ahead. \xB7\
- \ GitHub"
-- description: Click on Browser Use (on github browser-use at DuckDuckGo)
- output: null
- cssSelector: ''
- xpath: html/body/div[2]/div[6]/div[4]/div/div/div/div[2]/section[1]/ol/li[1]/article/div[3]/h2/a
- elementTag: a
- elementHash: 0a05391462
- selectorStrategies: null
- target_text: Browser Use
- container_hint: null
- position_hint: null
- interaction_type: null
- type: click
- agent_reasoning: 'The previous step successfully searched DuckDuckGo for "github
- browser-use" after hitting a rate limit on GitHub. The current page shows the
- search results. The immediate goal is to click the link that leads to the main
- "browser-use" repository on GitHub. The first result, `[1262]`, titled "GitHub
- - browser-use/browser-use: Make websites accessible for AI ...", with the URL
- `https://github.com/browser-use/browser-use`, is clearly the correct link. I will
- click the anchor element `[1219]` associated with this title.'
- page_context_url: https://duckduckgo.com/?q=github+browser-use&ia=web
- page_context_title: github browser-use at DuckDuckGo
-input_schema: []
diff --git a/workflows/workflow_use/healing/deterministic_converter.py b/workflows/workflow_use/healing/deterministic_converter.py
index 53e73a7a..9f1873e3 100644
--- a/workflows/workflow_use/healing/deterministic_converter.py
+++ b/workflows/workflow_use/healing/deterministic_converter.py
@@ -17,7 +17,8 @@ class DeterministicWorkflowConverter:
variable identification.
"""
- def __init__(self):
+ def __init__(self, llm=None):
+ self.llm = llm
self.element_text_map: Dict[str, str] = {} # Maps element hashes to visible text
self.element_hash_map: Dict[int, str] = {} # Maps element index to hash for selector population
self.captured_element_text_map: Dict[int, Any] = {} # Captured during agent execution
@@ -318,22 +319,23 @@ def _normalize_element_data(self, raw_data: Any) -> Dict[str, Any]:
return None
- def _extract_target_text(self, element_data: Optional[Dict[str, Any]], action_dict: Dict[str, Any]) -> str:
+ def _extract_target_text(
+ self, element_data: Optional[Dict[str, Any]], action_dict: Dict[str, Any], agent_context: Optional[Dict[str, Any]] = None
+ ) -> str:
"""
Extract the best target_text for semantic targeting from element data.
Priority:
1. Visible text content (node_value)
2. aria-label attribute
- 3. title attribute
- 4. placeholder attribute
+ 3. placeholder attribute (for input fields)
+ 4. title attribute
5. alt attribute (for images)
- 6. value attribute
- 7. name attribute
- 8. id attribute (last resort)
- 9. href attribute (for anchor tags) - extract meaningful part
- 10. Input text being entered (for input actions)
- 11. Node name + xpath hint (absolute fallback)
+ 6. name attribute (if human-readable, not technical ID)
+ 7. id attribute (only if human-readable, not technical ID)
+ 8. href attribute (for anchor tags) - extract meaningful part
+ 9. Input text being entered (for input actions)
+ 10. Node name + xpath hint (absolute fallback)
"""
if not element_data:
# For input actions, use the text being entered as fallback
@@ -341,23 +343,155 @@ def _extract_target_text(self, element_data: Optional[Dict[str, Any]], action_di
return action_dict['text']
return 'element'
- # Priority 1: Visible text content
+ # Priority 1: Visible text content (but NOT for input fields - they don't have meaningful text content)
+ node_name = element_data.get('node_name', '').lower()
node_value = element_data.get('node_value', '').strip()
- if node_value:
+
+ # Skip node_value for input fields (they don't have text content, only values)
+ if node_value and node_name not in ['input', 'textarea', 'select']:
print(f' ā Using node_value as target_text: "{node_value}"')
return node_value
- # Priority 2-8: Check attributes in order
+ # Priority 2-5: Check high-value attributes in order
attributes = element_data.get('attributes', {})
- for attr in ['aria-label', 'title', 'placeholder', 'alt', 'value', 'name', 'id']:
+ for attr in ['aria-label', 'placeholder', 'title', 'alt']:
if attr in attributes and attributes[attr]:
text = str(attributes[attr]).strip()
if text:
print(f' ā Using {attr} attribute as target_text: "{text}"')
return text
- # Priority 9: For anchor tags, extract meaningful text from href
- node_name = element_data.get('node_name', '')
+ # Priority 6: Extract from agent reasoning using structured [ELEMENT: "text"] format
+ # The agent is instructed to use this format: [ELEMENT: "First Name"], [ELEMENT: "Search"], etc.
+ if agent_context and agent_context.get('reasoning'):
+ reasoning = agent_context['reasoning']
+ import re
+
+ # Debug: print the reasoning to see what we're working with
+ print(f' š Agent reasoning: {reasoning[:200]}...')
+
+ # Primary Pattern: Extract from structured [ELEMENT: "text"] tag
+ # This is the most reliable since we explicitly ask the agent to use this format
+
+ # Find ALL [ELEMENT] tags and use the LAST one (closest to the action)
+ matches = list(re.finditer(r'\[ELEMENT:\s*["\']([^"\']+)["\']\]', reasoning))
+ if matches:
+ element_text = matches[-1].group(1).strip() # Use last match
+ print(f' ā Extracted from [ELEMENT] tag (last occurrence): "{element_text}"')
+ print(f' (Found {len(matches)} [ELEMENT] tags total, using the last one)')
+ return element_text
+
+ # Try without quotes as fallback: [ELEMENT: Search]
+ matches = list(re.finditer(r'\[ELEMENT:\s*([^\]]+)\]', reasoning))
+ if matches:
+ element_text = matches[-1].group(1).strip() # Use last match
+ print(f' ā Extracted from [ELEMENT] tag (no quotes, last occurrence): "{element_text}"')
+ return element_text
+
+ # Fallback patterns for when agent doesn't follow the structured format:
+
+ # For input/click actions, try to find context-specific field mentions
+ # E.g., "input 'Jasmine' into the First Name" or "'Paxton' into the Last Name"
+ action_value = action_dict.get('text') or action_dict.get('value')
+
+ if action_value:
+ # Pattern: Look for the value followed by field name mention
+ # E.g., "'Jasmine' into the First Name field" or "input 'Paxton'... Last Name"
+ escaped_value = re.escape(str(action_value))
+ # Match: value (with quotes or not) + optional words + "into/in/for" + field name + "field/input"
+ match = re.search(
+ rf'["\']?{escaped_value}["\']?[^.]*?(?:into|in|for|to)\s+(?:the\s+)?([A-Z][a-z]+(?:\s+[A-Z][a-z]+)*)\s+(?:field|input|box)',
+ reasoning,
+ re.IGNORECASE,
+ )
+ if match:
+ label_text = match.group(1).strip()
+ print(f' ā Extracted from agent reasoning (context: "{action_value}"): "{label_text}"')
+ return label_text
+
+ # Fallback: Pattern 1: "into the First Name field" (first occurrence)
+ match = re.search(r'(?:into|in|for)\s+(?:the\s+)?([A-Z][a-z]+(?:\s+[A-Z][a-z]+)*)\s+(?:field|input|box)', reasoning)
+ if match:
+ label_text = match.group(1).strip()
+ print(f' ā Extracted from agent reasoning: "{label_text}"')
+ return label_text
+
+ # Fallback: Pattern 2: "First Name field" or "Last Name input"
+ match = re.search(r'([A-Z][a-z]+(?:\s+[A-Z][a-z]+)*)\s+(?:field|input|box)', reasoning)
+ if match:
+ label_text = match.group(1).strip()
+ print(f' ā Extracted from agent reasoning: "{label_text}"')
+ return label_text
+
+ # Fallback: Pattern 3: "click the Search button" or "click on Search" (for button/link clicks)
+ match = re.search(
+ r'(?:click|tap|press)\s+(?:on\s+)?(?:the\s+)?([A-Z][a-z]+(?:\s+[A-Z][a-z]+)*)\s+(?:button|link)',
+ reasoning,
+ re.IGNORECASE,
+ )
+ if match:
+ button_text = match.group(1).strip()
+ print(f' ā Extracted button text from agent reasoning: "{button_text}"')
+ return button_text
+
+ # Priority 7-8: Check name/id attributes, but skip or convert technical/generated IDs
+ def is_human_readable(text: str) -> bool:
+ """Check if text is human-readable, not a technical ID."""
+ text_lower = text.lower()
+ # Skip if contains common technical patterns
+ technical_patterns = ['$', 'ctl', 'ctr', 'dnn', 'aspnet', 'viewstate', '__', 'guid']
+ if any(pattern in text_lower for pattern in technical_patterns):
+ return False
+ # Skip if mostly uppercase/numbers (like GUID fragments)
+ if len([c for c in text if c.isupper() or c.isdigit()]) > len(text) * 0.7:
+ return False
+ return True
+
+ def extract_semantic_part(technical_id: str) -> str | None:
+ """Try to extract semantic meaning from technical IDs like 'dnn$ctr434$SQLViewPro$FirstName$txtParameter'."""
+ # Split by common separators
+ parts = technical_id.replace('$', '.').replace('_', '.').split('.')
+
+ # Look for parts that might be semantic (e.g., "FirstName", "LastName", "Search")
+ for part in reversed(parts): # Check from end first (more specific)
+ # Skip common technical suffixes
+ if part.lower() in ['txt', 'txtparameter', 'parameter', 'ctrl', 'control', 'btn', 'button', 'lbl', 'label']:
+ continue
+ # Skip very short parts (likely not semantic)
+ if len(part) < 3:
+ continue
+ # Skip numeric parts
+ if part.isdigit():
+ continue
+ # Skip parts that look like prefixes (all caps)
+ if part.isupper() and len(part) < 5:
+ continue
+
+ # Found a potentially semantic part - convert camelCase to readable text
+ # E.g., "FirstName" -> "First Name"
+ import re
+
+ # Insert space before capital letters
+ readable = re.sub(r'([a-z])([A-Z])', r'\1 \2', part)
+ print(f' ā Extracted semantic text from {technical_id}: "{readable}"')
+ return readable
+
+ return None
+
+ for attr in ['name', 'id']:
+ if attr in attributes and attributes[attr]:
+ text = str(attributes[attr]).strip()
+ if text and is_human_readable(text):
+ print(f' ā Using {attr} attribute as target_text: "{text}"')
+ return text
+ elif text:
+ # Try to extract semantic meaning from technical IDs
+ semantic_text = extract_semantic_part(text)
+ if semantic_text:
+ return semantic_text
+ print(f' ā ļø Skipping technical {attr} attribute: "{text}"')
+
+ # Priority 8: For anchor tags, extract meaningful text from href
if node_name == 'a' and 'href' in attributes:
href = attributes['href']
if isinstance(href, str):
@@ -376,13 +510,8 @@ def _extract_target_text(self, element_data: Optional[Dict[str, Any]], action_di
print(f' ā Extracted from href as target_text: "{text}"')
return text
- # Priority 10: For input actions, use the text being entered
- if action_dict.get('text'):
- text = action_dict['text']
- print(f' ā Using input text as target_text: "{text}"')
- return text
-
- # Priority 11: Fallback - use node name as hint
+ # Priority 9: Fallback - use descriptive element type
+ # NEVER use action_dict.get('text') - that's the input VALUE, not a semantic identifier!
if node_name:
print(f' ā ļø No good target text found, using node name: "{node_name}"')
return f'{node_name} element'
@@ -433,7 +562,7 @@ def _convert_action_to_step(
# Input text actions (browser-use can use either 'input' or 'input_text')
elif action_type in ['input', 'input_text']:
- target_text = self._extract_target_text(element_data, action_dict)
+ target_text = self._extract_target_text(element_data, action_dict, agent_context)
# Ensure target_text is never empty
if not target_text:
target_text = 'input field'
@@ -463,11 +592,73 @@ def _convert_action_to_step(
# Click actions (browser-use uses 'click', not 'click_element')
elif action_type in ['click', 'click_element']:
- target_text = self._extract_target_text(element_data, action_dict)
+ target_text = self._extract_target_text(element_data, action_dict, agent_context)
# Ensure target_text is never empty
if not target_text:
target_text = 'element'
+ # Check if this looks like a dynamic identifier (ID, code, number, etc.) that should be made generic
+ import re
+
+ position_hint = None
+ container_hint = None
+
+ # Define common dynamic identifier patterns
+ # Require at least one digit to avoid matching regular words
+ alphanumeric_id = re.match(r'^[A-Z]{2,}\d{3,}$', target_text) # e.g., AP00945776, ABC123
+ numeric_id = re.match(r'^\d{3,}$', target_text) # e.g., 123456, 00945776
+ code_with_separator = re.match(
+ r'^[A-Z0-9]+[-_][A-Z0-9]*\d+[A-Z0-9]*$', target_text, re.IGNORECASE
+ ) # e.g., ORD-12345, user_456, TKT-9876
+
+ if alphanumeric_id or numeric_id or code_with_separator:
+ print(f' š Detected dynamic identifier pattern: "{target_text}"')
+
+ # Check agent reasoning for context to determine the semantic meaning
+ reasoning = agent_context.get('reasoning', '') if agent_context else ''
+ reasoning_lower = reasoning.lower()
+
+ original_target = target_text
+
+ # Map reasoning keywords to generic identifiers
+ # This makes workflows reusable across different records
+ semantic_mapping = {
+ 'license': 'license number link',
+ 'provider': 'provider id link',
+ 'order': 'order id link',
+ 'invoice': 'invoice number link',
+ 'ticket': 'ticket number link',
+ 'case': 'case number link',
+ 'patient': 'patient id link',
+ 'user': 'user id link',
+ 'customer': 'customer id link',
+ 'product': 'product code link',
+ 'transaction': 'transaction id link',
+ 'record': 'record id link',
+ }
+
+ # Try to find semantic meaning from reasoning
+ converted = False
+ for keyword, generic_name in semantic_mapping.items():
+ if keyword in reasoning_lower:
+ target_text = generic_name
+ position_hint = 'first' # Usually click the first result
+ container_hint = 'search results'
+ print(f' ā
Converted to generic: "{target_text}" (detected: {keyword})')
+ print(f' Position: {position_hint}, Container: {container_hint}')
+ print(f' Original value "{original_target}" will match via pattern')
+ converted = True
+ break
+
+ # If no semantic meaning found, use generic "id link"
+ if not converted:
+ target_text = 'id link'
+ position_hint = 'first'
+ container_hint = 'search results'
+ print(f' ā
Converted to generic: "{target_text}" (no specific context detected)')
+ print(f' Position: {position_hint}, Container: {container_hint}')
+ print(f' Original value "{original_target}" will match via pattern')
+
# Create semantic description
base_description = f'Click on {target_text}'
description = self._create_semantic_description(action_type, base_description, agent_context, target_text)
@@ -478,6 +669,12 @@ def _convert_action_to_step(
'description': description,
}
+ # Add position and container hints if detected
+ if position_hint:
+ step['position_hint'] = position_hint
+ if container_hint:
+ step['container_hint'] = container_hint
+
# Add element hash for selector population
if element_data and element_data.get('element_hash'):
step['elementHash'] = element_data['element_hash']
@@ -503,7 +700,7 @@ def _convert_action_to_step(
keys = action_dict.get('keys', '')
# Try to get target from last interacted element if available
- target_text = self._extract_target_text(element_data, action_dict)
+ target_text = self._extract_target_text(element_data, action_dict, agent_context)
# Ensure target_text is never empty
if not target_text:
target_text = 'page'
diff --git a/workflows/workflow_use/healing/service.py b/workflows/workflow_use/healing/service.py
index 9ff98c34..0adb0a67 100644
--- a/workflows/workflow_use/healing/service.py
+++ b/workflows/workflow_use/healing/service.py
@@ -27,7 +27,7 @@ def __init__(
self.enable_variable_extraction = enable_variable_extraction
self.use_deterministic_conversion = use_deterministic_conversion
self.variable_extractor = VariableExtractor(llm=llm) if enable_variable_extraction else None
- self.deterministic_converter = DeterministicWorkflowConverter() if use_deterministic_conversion else None
+ self.deterministic_converter = DeterministicWorkflowConverter(llm=llm) if use_deterministic_conversion else None
self.selector_generator = SelectorGenerator() # Initialize multi-strategy selector generator
self.interacted_elements_hash_map: dict[str, DOMInteractedElement] = {}
@@ -122,37 +122,15 @@ def _validate_workflow_quality(self, workflow_definition: WorkflowDefinitionSche
print()
def _populate_selector_fields(self, workflow_definition: WorkflowDefinitionSchema) -> WorkflowDefinitionSchema:
- """Populate cssSelector, xpath, and elementTag fields from interacted_elements_hash_map"""
- print('\nš§ Populating selector fields for workflow steps...')
+ """
+ DISABLED: We no longer populate xpath/cssSelector fields to rely purely on semantic matching.
+ This method is kept for backward compatibility but doesn't modify the workflow.
+ """
+ print('\nš§ Skipping selector field population - using pure semantic matching')
print(f' Available element hashes in map: {len(self.interacted_elements_hash_map)}')
- # Process each step to add back the selector fields
- populated_count = 0
- for i, step in enumerate(workflow_definition.steps):
- if isinstance(step, SelectorWorkflowSteps):
- print(f'\n Step {i + 1} (type={step.type}):')
- if hasattr(step, 'elementHash') and step.elementHash:
- print(f' elementHash: {step.elementHash}')
- if step.elementHash in self.interacted_elements_hash_map:
- dom_element = self.interacted_elements_hash_map[step.elementHash]
- # DOMInteractedElement has different attribute names
- step.cssSelector = getattr(dom_element, 'css_selector', '') or ''
- step.xpath = getattr(dom_element, 'x_path', '') or getattr(dom_element, 'xpath', '')
- step.elementTag = dom_element.node_name.lower() if hasattr(dom_element, 'node_name') else ''
-
- print(' ā
Populated:')
- print(f' cssSelector: {step.cssSelector[:80] if step.cssSelector else "(empty)"}')
- print(f' xpath: {step.xpath[:80] if step.xpath else "(empty)"}')
- print(f' elementTag: {step.elementTag}')
- populated_count += 1
- else:
- print(' ā ļø elementHash not found in map!')
- else:
- print(' (no elementHash)')
-
- print(f'\n ā
Populated {populated_count} steps with selector fields')
-
- # Create the full WorkflowDefinitionSchema with populated fields
+ # Just return the workflow as-is without populating xpath/cssSelector
+ # The semantic executor will use target_text for element matching
return workflow_definition
async def create_workflow_definition(
@@ -358,17 +336,24 @@ async def act(self, action, browser_session, *args, **kwargs):
tag_name = dom_element.get('tag_name', '')
attrs = dom_element.get('attributes', {})
else:
+ # Extract tag name first
+ tag_name = (
+ getattr(dom_element, 'node_name', '').lower() if hasattr(dom_element, 'node_name') else ''
+ )
+ attrs = getattr(dom_element, 'attributes', {})
+
# Extract text by trying multiple field names
text = ''
for text_field in ['text', 'inner_text', 'node_value', 'textContent', 'innerText']:
if hasattr(dom_element, text_field):
- text = getattr(dom_element, text_field, '')
- if text and text.strip():
+ potential_text = getattr(dom_element, text_field, '')
+ if potential_text and potential_text.strip():
+ # IMPORTANT: Skip JavaScript href text (same filter as in deterministic_converter.py)
+ # browser-use sometimes provides JavaScript href as 'text' for anchor tags
+ if tag_name == 'a' and potential_text.lower().startswith('javascript:'):
+ continue
+ text = potential_text
break
- tag_name = (
- getattr(dom_element, 'node_name', '').lower() if hasattr(dom_element, 'node_name') else ''
- )
- attrs = getattr(dom_element, 'attributes', {})
# Normalize text (strip whitespace)
text = text.strip() if text else ''
@@ -398,32 +383,61 @@ async def act(self, action, browser_session, *args, **kwargs):
if semantic_text:
text = semantic_text
print(f' š Using semantic attribute for better text: "{text}"')
- # For anchor tags with no good text, try parent text or href extraction
- elif tag_name == 'a' and 'href' in attrs:
- href = attrs['href']
- # Extract the last meaningful part of the URL path
- # E.g., "https://newsroom.edison.com/releases" -> "releases"
- if isinstance(href, str):
- # Remove query params and anchors
- href = href.split('?')[0].split('#')[0]
- # Get the last path segment
- path_parts = href.rstrip('/').split('/')
- if path_parts:
- last_part = path_parts[-1]
- # Only use if it looks like readable text
- # Avoid random IDs like "nboo9eyy" (all lowercase alphanumeric with no separators)
- if last_part and last_part not in ['www.edison.com', 'edison.com', 'investors']:
- # Check if it has word separators (hyphens, underscores)
- if '-' in last_part or '_' in last_part:
- text = last_part.replace('-', ' ').replace('_', ' ').title()
- print(f' š Extracted from href: "{text}"')
- # Fallback: use clean slugs without separators (e.g., "login", "dashboard")
- # Only if they're reasonable length and look like words (not random IDs)
- elif len(last_part) >= 3 and len(last_part) <= 20 and last_part.isalpha():
- text = last_part.title()
- print(f' š Extracted clean slug from href: "{text}"')
-
- # Final fallback for any element: if still no text, try attributes
+ # For anchor tags, try ID/class-based inference for common button patterns
+ elif tag_name == 'a':
+ element_id = attrs.get('id', '')
+ element_class = attrs.get('class', '')
+
+ # Check for common button patterns in ID/class
+ id_lower = element_id.lower() if element_id else ''
+ class_lower = element_class.lower() if element_class else ''
+
+ # Common search/submit button patterns
+ if 'search' in id_lower or 'search' in class_lower:
+ text = 'Search'
+ print(f' š Inferred "Search" from ID/class: {element_id or element_class}')
+ elif 'submit' in id_lower or 'submit' in class_lower:
+ text = 'Submit'
+ print(f' š Inferred "Submit" from ID/class: {element_id or element_class}')
+ elif 'action' in id_lower or 'action' in class_lower:
+ # cmdAction, btnAction, etc. in forms usually means Submit/Search
+ if 'sqlviewpro' in id_lower or 'parameter' in id_lower:
+ text = 'Search'
+ print(f' š Inferred "Search" from form action button: {element_id}')
+ else:
+ text = 'Submit'
+ print(f' š Inferred "Submit" from action button: {element_id}')
+ # If still no text after ID/class inference, try href extraction
+ elif 'href' in attrs:
+ href = attrs['href']
+ # Skip JavaScript hrefs - they don't have meaningful text to extract
+ if isinstance(href, str) and not href.lower().startswith('javascript:'):
+ # Extract the last meaningful part of the URL path
+ # E.g., "https://newsroom.edison.com/releases" -> "releases"
+ # Remove query params and anchors
+ href = href.split('?')[0].split('#')[0]
+ # Get the last path segment
+ path_parts = href.rstrip('/').split('/')
+ if path_parts:
+ last_part = path_parts[-1]
+ # Only use if it looks like readable text
+ # Avoid random IDs like "nboo9eyy" (all lowercase alphanumeric with no separators)
+ if last_part and last_part not in [
+ 'www.edison.com',
+ 'edison.com',
+ 'investors',
+ ]:
+ # Check if it has word separators (hyphens, underscores)
+ if '-' in last_part or '_' in last_part:
+ text = last_part.replace('-', ' ').replace('_', ' ').title()
+ print(f' š Extracted from href: "{text}"')
+ # Fallback: use clean slugs without separators (e.g., "login", "dashboard")
+ # Only if they're reasonable length and look like words (not random IDs)
+ elif len(last_part) >= 3 and len(last_part) <= 20 and last_part.isalpha():
+ text = last_part.title()
+ print(f' š Extracted clean slug from href: "{text}"')
+
+ # Final fallback for any element (not anchor/button): if still no text, try attributes
elif not text:
if isinstance(attrs, dict):
# Try common text attributes
@@ -435,6 +449,7 @@ async def act(self, action, browser_session, *args, **kwargs):
or attrs.get('value')
or ''
)
+ # Note: ID/class inference for anchor tags is now handled above in the anchor/button block
# Create a simplified dict with the data we need
# Handle both dict and object formats
@@ -480,13 +495,32 @@ async def act(self, action, browser_session, *args, **kwargs):
result = await super().act(action, browser_session, *args, **kwargs)
return result
+ # Enhance the prompt to ensure agent mentions visible text of elements in a structured format
+ enhanced_prompt = f"""{prompt}
+
+CRITICAL WORKFLOW GENERATION REQUIREMENTS:
+For EVERY action you take, you MUST include this structured tag in your reasoning:
+
+Format: [ELEMENT: "exact visible text"]
+
+Examples:
+- "I will input 'John' [ELEMENT: "First Name"] into the form"
+- "I will input 'Doe' [ELEMENT: "Last Name"] into the form"
+- "I will click [ELEMENT: "Search"] to submit the form"
+- "I will click [ELEMENT: "License Number"] to view details"
+- "I will select [ELEMENT: "Country"] from the dropdown"
+
+The [ELEMENT: "..."] tag must contain the EXACT visible text of the button, label, link, or field you're interacting with.
+This structured format is critical for generating a reusable workflow."""
+
agent = Agent(
- task=prompt,
+ task=enhanced_prompt,
browser_session=browser,
llm=agent_llm,
page_extraction_llm=extraction_llm,
controller=CapturingController(self.selector_generator), # Pass selector_generator to controller
enable_memory=False,
+ use_vision=True,
max_failures=10,
)
diff --git a/workflows/workflow_use/storage/service.py b/workflows/workflow_use/storage/service.py
index 860724f8..9972a934 100644
--- a/workflows/workflow_use/storage/service.py
+++ b/workflows/workflow_use/storage/service.py
@@ -123,8 +123,55 @@ def save_workflow(
logger.info(f'Creating new workflow: {workflow_id}')
# Save workflow file
+ # Exclude legacy/unnecessary fields to keep workflow files clean
+ workflow_dict = workflow.model_dump(mode='json', exclude_none=True)
+
+ # Remove top-level bloat
+ workflow_dict.pop('workflow_analysis', None)
+
+ # Clean up steps by removing legacy/verbose/null fields
+ if 'steps' in workflow_dict:
+ for step in workflow_dict['steps']:
+ # Remove legacy/verbose fields (but keep xpath as fallback selector)
+ fields_to_remove = [
+ 'agent_reasoning', # Verbose agent thinking
+ 'page_context_url', # Redundant context
+ 'page_context_title', # Redundant context
+ 'elementHash', # Internal hash
+ 'elementTag', # Can be inferred
+ 'container_hint', # Usually null
+ 'position_hint', # Usually null
+ 'interaction_type', # Usually null
+ 'default_value', # Usually null for inputs
+ 'output', # Usually null
+ ]
+ for field in fields_to_remove:
+ step.pop(field, None)
+
+ # Remove empty cssSelector
+ if step.get('cssSelector') == '':
+ step.pop('cssSelector', None)
+
+ # Clean up selectorStrategies - remove if null or if all values are JavaScript href
+ if 'selectorStrategies' in step:
+ strategies = step['selectorStrategies']
+ if strategies is None:
+ step.pop('selectorStrategies', None)
+ elif isinstance(strategies, list):
+ # Filter out strategies with JavaScript href values
+ cleaned_strategies = [
+ s
+ for s in strategies
+ if not (isinstance(s.get('value'), str) and s.get('value', '').lower().startswith('javascript:'))
+ ]
+ # Only keep if we have valid strategies
+ if cleaned_strategies:
+ step['selectorStrategies'] = cleaned_strategies
+ else:
+ step.pop('selectorStrategies', None)
+
with open(metadata.file_path, 'w') as f:
- yaml.dump(workflow.model_dump(mode='json'), f, default_flow_style=False, sort_keys=False)
+ yaml.dump(workflow_dict, f, default_flow_style=False, sort_keys=False)
# Update metadata
self.metadata[workflow_id] = metadata
diff --git a/workflows/workflow_use/workflow/semantic_executor.py b/workflows/workflow_use/workflow/semantic_executor.py
index 0f6e6849..67bfb6f5 100644
--- a/workflows/workflow_use/workflow/semantic_executor.py
+++ b/workflows/workflow_use/workflow/semantic_executor.py
@@ -289,16 +289,19 @@ def _find_element_by_text(self, target_text: str, context_hints: List[str] = Non
)
return best_hierarchical_match
- # Strategy 2: Try partial matches with different strategies (original fallback)
+ # Strategy 2: Try partial matches with different strategies (including label_text for input fields)
for text, element_info in self.current_mapping.items():
text_lower = text.lower()
original_text = element_info.get('original_text', '').lower()
+ # IMPORTANT: Check label_text for input fields (labels are in separate elements)
+ label_text = element_info.get('label_text', '').lower()
- # Check if target text is contained in element text (more lenient)
+ # Check if target text is contained in element text, original text, OR label text
if (
target_lower in text_lower
or text_lower in target_lower
or (original_text and (target_lower in original_text or original_text in target_lower))
+ or (label_text and (target_lower in label_text or label_text in target_lower))
):
# For radio buttons and checkboxes, be more specific
if element_info.get('element_type') in ['radio', 'checkbox']:
@@ -343,6 +346,176 @@ def _find_element_by_text(self, target_text: str, context_hints: List[str] = Non
return None
+ def _find_element_by_pattern(
+ self, pattern: str, position_hint: Optional[str] = None, container_hint: Optional[str] = None
+ ) -> Optional[Dict]:
+ """
+ Find element by pattern matching for dynamic identifiers using priority-based strategies.
+ This is a generic method that works for any dynamic content (IDs, codes, numbers, etc.)
+
+ Args:
+ pattern: The pattern text to match (e.g., "license number link", "order id", "product code")
+ position_hint: Position hint like "first", "last", "second", or numeric index
+ container_hint: Container context like "search results", "table", "list"
+
+ Returns:
+ Element info dict if found, None otherwise
+ """
+ import re
+
+ logger.info(f"Finding element by pattern: '{pattern}' (position: {position_hint}, container: {container_hint})")
+
+ # Define common dynamic identifier patterns (alphanumeric codes/IDs)
+ # Require at least one digit to avoid matching regular words
+ alphanumeric_id_pattern = re.compile(r'^[A-Z]{2,}\d{3,}$') # e.g., AP00945776, ABC123, XY12345
+ numeric_id_pattern = re.compile(r'^\d{3,}$') # e.g., 123456, 00945776
+ code_pattern = re.compile(r'^[A-Z0-9]+[-_][A-Z0-9]*\d+[A-Z0-9]*$', re.IGNORECASE) # e.g., ORD-12345, user_456, TKT-9876
+
+ matching_elements = []
+
+ # Priority 1: Exact text match in semantic mapping (highest priority)
+ # This handles cases where the exact dynamic value was captured
+ logger.debug(f'[Priority 1] Searching for exact text matches')
+ for text, element_info in self.current_mapping.items():
+ text_stripped = text.split(' (in ')[0].strip() # Remove context annotations
+
+ # Check if text matches common ID/code patterns
+ if (
+ alphanumeric_id_pattern.match(text_stripped)
+ or numeric_id_pattern.match(text_stripped)
+ or code_pattern.match(text_stripped)
+ ):
+ matching_elements.append((text, element_info, 1))
+ logger.debug(f'[Priority 1] Found ID/code pattern: {text_stripped}')
+
+ if matching_elements:
+ logger.info(f'ā
Found {len(matching_elements)} exact ID/code matches (Priority 1)')
+ return self._select_element_by_position(matching_elements, position_hint, container_hint)
+
+ # Priority 2: Clickable elements in structured containers (tables, lists) with ID-like patterns
+ # This is common for search results, data grids, etc.
+ logger.info('No exact matches, trying Priority 2: Clickable elements in structured containers')
+ matching_elements = []
+ structured_containers = ['table', 'cell', 'td', 'tr', 'list', 'ul', 'ol', 'li', 'grid', 'row']
+
+ for text, element_info in self.current_mapping.items():
+ text_stripped = text.split(' (in ')[0].strip()
+ context = text.split(' (in ')[1].rstrip(')') if ' (in ' in text else ''
+
+ # Check if element is in a structured container
+ in_structured_container = any(container in context.lower() for container in structured_containers)
+
+ if in_structured_container:
+ element_tag = element_info.get('element_type', '')
+ # Look for clickable elements (links, buttons)
+ if element_tag in ['a', 'link', 'button']:
+ # Check if text looks like an ID/code
+ if (
+ alphanumeric_id_pattern.match(text_stripped)
+ or numeric_id_pattern.match(text_stripped)
+ or code_pattern.match(text_stripped)
+ ):
+ matching_elements.append((text, element_info, 2))
+ logger.debug(f'[Priority 2] Found clickable ID in {context}: {text_stripped}')
+
+ if matching_elements:
+ logger.info(f'ā
Found {len(matching_elements)} clickable IDs in structured containers (Priority 2)')
+ return self._select_element_by_position(matching_elements, position_hint, container_hint)
+
+ # Priority 3: Fuzzy match using pattern keywords
+ # Extract meaningful keywords from the pattern (e.g., "license number link" -> ["license", "number", "link"])
+ logger.info('No structured matches, trying Priority 3: Fuzzy keyword matching')
+ matching_elements = []
+ pattern_keywords = [word for word in pattern.lower().split() if len(word) > 3]
+
+ for text, element_info in self.current_mapping.items():
+ text_stripped = text.split(' (in ')[0].strip()
+ element_tag = element_info.get('element_type', '')
+
+ # Look for clickable elements
+ if element_tag in ['a', 'link', 'button']:
+ text_lower = text.lower()
+ # Check if any significant keyword appears
+ keyword_matches = sum(1 for keyword in pattern_keywords if keyword in text_lower)
+ if keyword_matches > 0:
+ matching_elements.append((text, element_info, 3, keyword_matches)) # Include match count for sorting
+ logger.debug(f'[Priority 3] Found {keyword_matches} keyword match(es): {text_stripped}')
+
+ if matching_elements:
+ # Sort by number of keyword matches (descending)
+ matching_elements.sort(key=lambda x: x[3], reverse=True)
+ # Convert back to (text, element_info, priority) format
+ matching_elements = [(t, e, p) for t, e, p, _ in matching_elements]
+ logger.info(f'ā
Found {len(matching_elements)} keyword matches (Priority 3)')
+ return self._select_element_by_position(matching_elements, position_hint, container_hint)
+
+ # Priority 4: Container-based selection (lowest priority)
+ # Use container_hint to narrow down to a specific region, then select by position
+ if container_hint:
+ logger.info(f"No keyword matches, trying Priority 4: Any clickable in '{container_hint}' container")
+ matching_elements = []
+
+ for text, element_info in self.current_mapping.items():
+ context = text.split(' (in ')[1].rstrip(')') if ' (in ' in text else ''
+ element_tag = element_info.get('element_type', '')
+
+ # Match by container hint in context
+ if container_hint.lower() in context.lower():
+ if element_tag in ['a', 'link', 'button']:
+ matching_elements.append((text, element_info, 4))
+ logger.debug(f'[Priority 4] Found clickable in {container_hint}: {text.split(" (in ")[0].strip()}')
+
+ if matching_elements:
+ logger.info(f'ā
Found {len(matching_elements)} clickable elements in container (Priority 4)')
+ return self._select_element_by_position(matching_elements, position_hint, container_hint)
+
+ # No matches found at any priority level
+ logger.warning(f"ā No elements found matching pattern '{pattern}' at any priority level")
+ return None
+
+ def _select_element_by_position(
+ self, matching_elements: list, position_hint: Optional[str], container_hint: Optional[str]
+ ) -> Optional[Dict]:
+ """
+ Select element from matching_elements based on position hint.
+ matching_elements is a list of tuples: (text, element_info, priority)
+ """
+ if not matching_elements:
+ return None
+
+ # Sort by priority (lower number = higher priority)
+ matching_elements.sort(key=lambda x: x[2])
+
+ # Apply position hint
+ if position_hint == 'first' and matching_elements:
+ selected_text, selected_element, priority = matching_elements[0]
+ logger.info(f'Selected first matching element (Priority {priority}): {selected_text}')
+ return selected_element
+ elif position_hint == 'last' and matching_elements:
+ # Get all elements with the best priority
+ best_priority = matching_elements[0][2]
+ best_matches = [e for e in matching_elements if e[2] == best_priority]
+ selected_text, selected_element, priority = best_matches[-1]
+ logger.info(f'Selected last matching element (Priority {priority}): {selected_text}')
+ return selected_element
+ elif position_hint and position_hint.isdigit():
+ index = int(position_hint) - 1 # Convert to 0-indexed
+ # Get all elements with the best priority
+ best_priority = matching_elements[0][2]
+ best_matches = [e for e in matching_elements if e[2] == best_priority]
+ if 0 <= index < len(best_matches):
+ selected_text, selected_element, priority = best_matches[index]
+ logger.info(f'Selected element at position {position_hint} (Priority {priority}): {selected_text}')
+ return selected_element
+
+ # No position hint or invalid position - return first match (highest priority)
+ if matching_elements:
+ selected_text, selected_element, priority = matching_elements[0]
+ logger.info(f'No valid position hint, returning first match (Priority {priority}): {selected_text}')
+ return selected_element
+
+ return None
+
async def _try_direct_selector(self, target_text: str) -> Optional[str]:
"""Try to use target_text as a direct selector (ID or name) with improved robustness."""
if not target_text or not target_text.replace('_', '').replace('-', '').replace('.', '').isalnum():
@@ -572,8 +745,18 @@ async def execute_click_step(self, step: ClickStep) -> ActionResult:
if hasattr(step, 'target_text') and step.target_text:
target_identifier = step.target_text
+ # Check for position and container hints for dynamic elements
+ position_hint = getattr(step, 'position_hint', None)
+ container_hint = getattr(step, 'container_hint', None)
+
# Always try to get element_info from semantic mapping for metadata (element_type, etc.)
- element_info = self._find_element_by_text(step.target_text)
+ # If we have hints, use them for more flexible matching
+ if position_hint or container_hint:
+ logger.info(f'Using hints - position: {position_hint}, container: {container_hint}')
+ # Find elements by pattern (e.g., "license number link" matches any license number)
+ element_info = self._find_element_by_pattern(step.target_text, position_hint, container_hint)
+ else:
+ element_info = self._find_element_by_text(step.target_text)
# SEMANTIC WORKFLOW PRIORITY:
# 1. Try direct selector by ID/name (stable, semantic attributes)
@@ -944,7 +1127,8 @@ async def _click_element_intelligently(self, selector: str, target_text: str, el
# Check if this is a submit button that should trigger navigation
is_submit_button = button_type == 'submit' or any(
- keyword in target_text.lower() for keyword in ['next', 'submit', 'continue', 'save', 'finish']
+ keyword in target_text.lower()
+ for keyword in ['next', 'submit', 'continue', 'save', 'finish', 'search']
)
if is_submit_button:
@@ -970,44 +1154,23 @@ async def _click_element_intelligently(self, selector: str, target_text: str, el
logger.info(f'ā
Navigation successful: {current_url} -> {new_url}')
return True
else:
- logger.warning(
- f"ā ļø No navigation occurred after clicking '{target_text}'. URL still: {new_url}"
+ # URL didn't change - could be same-page postback (ASP.NET) or AJAX update
+ logger.info(
+ f"ā ļø URL didn't change after clicking '{target_text}' (URL: {new_url}). Checking for page updates..."
)
- # Check for validation errors
+ # Wait a bit more for dynamic content to load
+ await asyncio.sleep(1)
+
+ # For same-page updates, assume success if no validation errors
+ # The semantic mapping will be refreshed before the next step
validation_errors = await self._detect_form_validation_errors()
if validation_errors:
logger.error(f'ā Form validation errors preventing submission: {validation_errors}')
return False
- # No validation errors but still no navigation - try submitting form directly
- logger.info('š§ Attempting to submit parent form directly via JavaScript')
- try:
- # Find parent form and submit it
- form_submitted = await self._element_evaluate(
- button_element,
- '(function() { var form = this.closest("form"); if(form) { form.requestSubmit(); return true; } return false; })',
- )
-
- if form_submitted:
- logger.info('ā
Triggered form.requestSubmit()')
- await asyncio.sleep(2)
- final_url = await page.get_url()
-
- if final_url != current_url:
- logger.info(
- f'ā
Navigation successful via form submit: {current_url} -> {final_url}'
- )
- return True
- else:
- logger.error("ā Form submit didn't trigger navigation either")
- return False
- else:
- logger.warning('ā ļø Could not find parent form to submit')
- return False
- except Exception as submit_error:
- logger.error(f'ā Error submitting form directly: {submit_error}')
- return False
+ logger.info('ā
No validation errors detected - assuming same-page update succeeded')
+ return True
except Exception as e:
logger.error(f'ā Error during submit button handling: {e}')
diff --git a/workflows/workflow_use/workflow/semantic_extractor.py b/workflows/workflow_use/workflow/semantic_extractor.py
index 644acba8..de36baef 100644
--- a/workflows/workflow_use/workflow/semantic_extractor.py
+++ b/workflows/workflow_use/workflow/semantic_extractor.py
@@ -536,7 +536,22 @@ async def extract_interactive_elements(self, page: 'Page') -> List[Dict]:
labelText = prevElement.textContent?.trim() || '';
}
}
-
+
+ // IMPORTANT: Handle table structures where label is in previous
+ if (!labelText) {
+ const parentCell = el.closest('td, th');
+ if (parentCell) {
+ const prevCell = parentCell.previousElementSibling;
+ if (prevCell && (prevCell.tagName === 'TD' || prevCell.tagName === 'TH')) {
+ const cellText = prevCell.textContent?.trim() || '';
+ // Only use if it looks like a label (short text, ends with colon, etc.)
+ if (cellText && cellText.length < 50) {
+ labelText = cellText.replace(/[:ļ¼]\s*$/, '').trim();
+ }
+ }
+ }
+ }
+
return labelText;
} catch (error) {
debugMessage('Error in safeGetLabelText', error.message);
@@ -838,6 +853,7 @@ async def extract_semantic_mapping(self, page: 'Page') -> Dict[str, Dict]:
'element_type': element_type,
'deterministic_id': element_id,
'original_text': text,
+ 'label_text': element_info.get('label_text', ''), # IMPORTANT: Include label text for input field matching
'dom_path': element_info.get('dom_path', ''),
'container_context': element_info.get('container_context', {}),
'sibling_context': element_info.get('sibling_context', {}),
diff --git a/workflows/workflow_use/workflow/service.py b/workflows/workflow_use/workflow/service.py
index 3b70c572..284e06c0 100644
--- a/workflows/workflow_use/workflow/service.py
+++ b/workflows/workflow_use/workflow/service.py
@@ -49,6 +49,9 @@ def __init__(
page_extraction_llm: BaseChatModel | None = None,
fallback_to_agent: bool = True,
use_cloud: bool = False,
+ debug: bool = False,
+ debug_log_folder: str | Path | None = None,
+ step_wait_time: float = 0.1,
) -> None:
"""Initialize a new Workflow instance from a schema object.
@@ -59,6 +62,9 @@ def __init__(
llm: Optional language model for fallback agent functionality
fallback_to_agent: Whether to fall back to agent-based execution on step failure
use_cloud: Whether to use browser-use cloud browser service instead of local browser
+ debug: Whether to enable debug mode (captures screenshots for each step)
+ debug_log_folder: Custom folder path for debug logs and screenshots (default: ./logs/workflow_debug)
+ step_wait_time: Time to wait between steps in seconds (default: 0.1)
Raises:
ValueError: If the workflow schema is invalid (though Pydantic handles most).
@@ -77,6 +83,13 @@ def __init__(
self.fallback_to_agent = fallback_to_agent
+ # Debug mode settings
+ self.debug = debug
+ self.debug_log_folder = Path(debug_log_folder) if debug_log_folder else Path('./logs/workflow_debug')
+
+ # Step execution settings
+ self.step_wait_time = step_wait_time
+
# Initialize multi-strategy element finder
self.element_finder = ElementFinder()
@@ -96,6 +109,9 @@ def load_from_file(
browser: Browser | None = None,
page_extraction_llm: BaseChatModel | None = None,
use_cloud: bool = False,
+ debug: bool = False,
+ debug_log_folder: str | Path | None = None,
+ step_wait_time: float = 0.1,
) -> Workflow:
"""Load a workflow from a file."""
with open(file_path, 'r', encoding='utf-8') as f:
@@ -108,6 +124,9 @@ def load_from_file(
llm=llm,
page_extraction_llm=page_extraction_llm,
use_cloud=use_cloud,
+ debug=debug,
+ debug_log_folder=debug_log_folder,
+ step_wait_time=step_wait_time,
)
# --- Runners ---
@@ -297,6 +316,73 @@ async def _run_agent_step(self, step: AgenticWorkflowStep, step_index: int) -> A
return await agent.run()
+ async def _run_extraction_step(self, step, step_index: int) -> ActionResult:
+ """
+ Lightweight extraction that uses LLM directly without spinning up an agent.
+ Much faster and cheaper than using a full agent for simple page content extraction.
+ """
+ from browser_use.agent.views import ActionResult
+
+ # Get extraction goal
+ extraction_goal = ''
+ if hasattr(step, 'goal'):
+ extraction_goal = step.goal
+ elif hasattr(step, 'extractionGoal'):
+ extraction_goal = step.extractionGoal
+ else:
+ extraction_goal = 'Extract information from the page'
+
+ # Get current page content using markdown extraction
+ page = await self.browser.get_current_page()
+ page_text, _ = await page._extract_clean_markdown()
+ page_url = await page.get_url()
+
+ # Use page_extraction_llm if available, otherwise fall back to main llm
+ extraction_llm = self.page_extraction_llm or self.llm
+
+ # Build extraction prompt
+ # Limit page text to avoid token limits
+ truncated_page_text = page_text[:10000] if page_text else ''
+
+ extraction_prompt = f"""You are extracting information from a web page.
+
+Page URL: {page_url}
+
+Extraction Goal:
+{extraction_goal}
+
+Page Content:
+{truncated_page_text}
+
+Instructions:
+- Extract the requested information accurately
+- Return ONLY the extracted data, no explanations
+- If the information is not found, return an empty string or appropriate null value
+- Format the output as requested in the extraction goal
+
+Extracted Information:"""
+
+ # Call LLM directly
+ messages = [UserMessage(content=extraction_prompt)]
+ response = await extraction_llm.ainvoke(messages)
+
+ # Extract the text content from response
+ # ainvoke returns ChatInvokeCompletion with a 'completion' attribute
+ extracted_content = ''
+ if hasattr(response, 'completion'):
+ extracted_content = response.completion
+ elif isinstance(response, str):
+ extracted_content = response
+
+ logger.info(f'Extracted content: {extracted_content[:200]}...')
+
+ # Return as ActionResult
+ return ActionResult(
+ is_done=False,
+ extracted_content=extracted_content,
+ include_in_memory=True,
+ )
+
# async def _fallback_to_agent(
# self,
# step_resolved: WorkflowStep,
@@ -516,27 +602,11 @@ async def _execute_step(self, step_index: int, step_resolved: WorkflowStep) -> A
action_name = step_resolved.type or '[No action specified]'
- # Extraction steps ALWAYS use agent/LLM, even in deterministic mode
+ # Extraction steps use lightweight LLM extraction (no agent needed)
is_extraction_step = action_name in ['extract', 'extract_page_content']
if is_extraction_step:
- logger.info(f'Step {step_index + 1}: Extraction step detected - using agent with LLM')
- # Convert to agent step
- from workflow_use.schema.views import AgentTaskWorkflowStep
-
- # Create agent task based on extraction step type
- if action_name == 'extract':
- task = getattr(step_resolved, 'extractionGoal', 'Extract information from the page')
- else: # extract_page_content
- task = getattr(step_resolved, 'goal', 'Extract text from the page')
-
- # Create an AgentTaskWorkflowStep on the fly
- agent_step = AgentTaskWorkflowStep(
- type='agent',
- task=task,
- description=step_resolved.description if hasattr(step_resolved, 'description') else None,
- output=step_resolved.output if hasattr(step_resolved, 'output') else None,
- )
- result = await self._run_agent_step(agent_step, step_index)
+ logger.info(f'Step {step_index + 1}: Extraction step detected - using lightweight LLM extraction')
+ result = await self._run_extraction_step(step_resolved, step_index)
# Check if this is a selector step without cssSelector - use semantic execution
elif action_name in ['click', 'input', 'key_press', 'select_change']:
@@ -702,6 +772,47 @@ async def run_step(self, step_index: int, inputs: dict[str, Any] | None = None):
# await self.browser.close() # <-- Commented out for testing
return result
+ async def _capture_debug_screenshot(self, step_index: int, step_description: str, prefix: str = '') -> None:
+ """Capture a screenshot for debugging purposes.
+
+ Args:
+ step_index: The index of the current step
+ step_description: Description of the step for the filename
+ prefix: Optional prefix for the filename (e.g., 'before', 'after', 'error')
+ """
+ if not self.debug:
+ return
+
+ try:
+ # Create debug log folder if it doesn't exist
+ self.debug_log_folder.mkdir(parents=True, exist_ok=True)
+
+ # Clean step description for filename (remove special characters)
+ import re
+
+ clean_description = re.sub(r'[^\w\s-]', '', step_description)
+ clean_description = re.sub(r'[-\s]+', '_', clean_description)
+ clean_description = clean_description[:50] # Limit length
+
+ # Create timestamp for uniqueness
+ from datetime import datetime
+
+ timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
+
+ # Build filename
+ prefix_str = f'{prefix}_' if prefix else ''
+ filename = f'step_{step_index + 1:02d}_{prefix_str}{clean_description}_{timestamp}.png'
+ screenshot_path = self.debug_log_folder / filename
+
+ # Capture screenshot
+ page = await self.browser.get_current_page()
+ await page.screenshot(path=str(screenshot_path), full_page=True)
+
+ logger.info(f'šø Debug screenshot saved: {screenshot_path}')
+
+ except Exception as e:
+ logger.warning(f'Failed to capture debug screenshot: {e}')
+
async def run(
self,
inputs: dict[str, Any] | None = None,
@@ -730,10 +841,24 @@ async def run(
results: List[ActionResult | AgentHistoryList] = []
+ # Log debug mode status
+ if self.debug:
+ logger.info(f'š Debug mode enabled - screenshots will be saved to: {self.debug_log_folder}')
+ # Create debug folder at the start
+ self.debug_log_folder.mkdir(parents=True, exist_ok=True)
+
+ # Log step wait time if configured
+ if self.step_wait_time > 0.1: # Only log if it's been changed from default
+ logger.info(f'ā±ļø Step wait time configured: {self.step_wait_time}s between steps')
+
await self.browser.start()
try:
for step_index, step_dict in enumerate(self.schema.steps): # self.steps now holds dictionaries
- await asyncio.sleep(0.1)
+ # Wait between steps (configurable)
+ if step_index > 0: # Don't wait before the first step
+ await asyncio.sleep(self.step_wait_time)
+ if self.step_wait_time > 0:
+ logger.debug(f'Waited {self.step_wait_time}s between steps')
# Check if cancellation was requested
if cancel_event and cancel_event.is_set():
@@ -743,11 +868,23 @@ async def run(
# Use description from the step dictionary
step_description = step_dict.description or 'No description provided'
logger.info(f'--- Running Step {step_index + 1}/{len(self.schema.steps)} -- {step_description} ---')
+
+ # Capture screenshot before step execution (if debug enabled)
+ await self._capture_debug_screenshot(step_index, step_description, prefix='before')
+
# Resolve placeholders using the current context (works on the dictionary)
step_resolved = self._resolve_placeholders(step_dict)
# Execute step using the unified _execute_step method
- result = await self._execute_step(step_index, step_resolved)
+ try:
+ result = await self._execute_step(step_index, step_resolved)
+
+ # Capture screenshot after successful step execution (if debug enabled)
+ await self._capture_debug_screenshot(step_index, step_description, prefix='after')
+ except Exception as e:
+ # Capture screenshot on error (if debug enabled)
+ await self._capture_debug_screenshot(step_index, step_description, prefix='error')
+ raise # Re-raise the exception after capturing screenshot
results.append(result)
# Persist outputs using the resolved step dictionary