Scale AI Software Engineer Interview Questions
14+ questions from real Scale AI Software Engineer interviews, reported by candidates.
Round Types
Top Topics
Questions
Scale.ai Enterprise GenAI SDE Fulltime Tech Phone Screen Experience
Fresh interview experience. I interviewed for Enterprise Genai SDE Phone Screen. It was a standard LLD (Legal Management) role. The questions were similar to: given an event and some JSON data, write
Overall, it went pretty well. I failed the backend practical section, and the questions seemed to be identical to those on the forum. The main thing is to learn the LLM API and how to use it. Not much
Round 1 — BQ + Backend Practical Behavioral questions covered previous projects. For the backend practical the stack chosen was Python. The task was similar to implementing a lightweight load balancer
## Problem Design an object-oriented engine for a simplified turn-based card game. Two players take turns drawing and playing cards. Cards have types: `Attack` (reduces opponent HP), `Heal` (restores own HP), `Shield` (blocks next attack). Game ends when a player's HP reaches 0. Design the class structure, then implement `play_turn(player, card)`. ```python class Card: def __init__(self, name: str, card_type: str, value: int): ... class Player: def __init__(self, name: str, hp: int, deck: list): ... def draw(self) -> Card: ... def is_alive(self) -> bool: ... class Game: def __init__(self, player_a: Player, player_b: Player): ... def play_turn(self, attacker: Player, card: Card) -> str: ... # returns a description of what happened def is_over(self) -> bool: ... def winner(self): ... ``` **Example turn:** ``` card = Card("Fireball", "Attack", 20) game.play_turn(alice, card) # "Alice plays Fireball. Bob takes 20 damage. Bob HP: 80." ``` ## Follow-ups 1. How would you implement the `Shield` mechanic so it persists across turns and stacks correctly? 2. How would you add a `Deck` class with shuffle, draw, and discard-pile management? 3. How would you extend the engine to support card effects that trigger on specific game events (e.g., "on draw" or "on death")? 4. How would you serialize the full game state to JSON for save/resume functionality?
**Format:** Practical onsite round — mix of live debugging, REST API design, and SQL query writing. Expect a shared editor with pre-written code containing bugs. **Section 1 — Debugging** You are given a Flask endpoint that is returning 500 errors intermittently: ```python @app.route("/orders/<order_id>") def get_order(order_id): conn = db.connect() # Bug: connection never closed row = conn.execute( f"SELECT * FROM orders WHERE id = {order_id}" # Bug: SQL injection ).fetchone() return jsonify(row) # Bug: row is None when not found -> TypeError ``` Identify and fix all three bugs. **Section 2 — API Design** Design REST endpoints for a basic order management system: create order, update order status, list orders by customer with pagination. Discuss: idempotency keys, HTTP status code choices, versioning strategy. **Section 3 — SQL** Given `orders(id, customer_id, status, created_at)` and `order_items(order_id, product_id, quantity, unit_price)`, write a query returning the top 5 customers by total spend in the last 30 days. ## Follow-ups 1. How would you add rate limiting to the orders endpoint without modifying application code? 2. What index would you add to optimize the top-customers query? 3. How would you test the SQL injection fix to ensure it is actually safe?
**Format:** Live debugging round — you are given a codebase with several interconnected bugs across a data pipeline. You must diagnose each bug, explain the root cause, and apply a fix. **Bug 1 — Off-by-one in pagination:** ```python def paginate(items, page, page_size): start = page * page_size # Bug: page=1 should start at index page_size return items[start:start+page_size] # Fix: start = (page - 1) * page_size ``` **Bug 2 — Mutable default argument:** ```python def append_event(event, history=[]): # Bug: shared across all calls history.append(event) return history # Fix: def append_event(event, history=None): history = history or [] ``` **Bug 3 — Silent exception swallowing:** ```python try: result = process(data) except Exception: pass # Bug: hides all errors, downstream uses stale result # Fix: at minimum log the exception; re-raise or set result=None and handle ``` **Bug 4 — Race condition in cache update:** ```python if key not in cache: cache[key] = fetch(key) # Bug: two threads can both enter this branch # Fix: use a lock or atomic setdefault / cache library with locking ``` ## Follow-ups 1. How would you write a regression test for the mutable default argument bug? 2. What linting tools catch these categories of bugs automatically in Python? 3. How would you set up structured logging so silent exception swallowing is impossible by convention? 4. What is the difference between a race condition and a deadlock? Give an example of each.
## Round 1 - Coding ## Problem Design and implement a simplified card game engine. You are given a standard 52-card deck. Players take turns drawing and playing cards. Implement the core game loop. ```python class Card: def __init__(self, suit: str, rank: str): ... class Deck: def shuffle(self) -> None: ... def draw(self) -> Card: ... class Player: def __init__(self, name: str): ... def play_card(self, index: int) -> Card: ... class CardGame: def __init__(self, players: list[Player]): ... def start(self) -> None: ... def get_winner(self) -> Player: ... ``` ``` Example: game = CardGame([Player("Alice"), Player("Bob")]) game.start() # Each player draws 5 cards; highest-rank card wins each round. # After 5 rounds: winner = player with most round wins game.get_winner() -> Player("Alice") ``` ## Follow-ups 1. How would you handle ties when two players play the same rank? 2. Extend to support 4 players - what changes in your design? 3. How would you make the game rules configurable (e.g. swap rule engine)? 4. If game state must be persisted mid-round, how do you serialize it?
## Round 1 - Coding ## Problem You are given a directed acyclic graph representing a neural network. Each node (neuron) has an initial activation state of 0 or 1. A neuron fires (state = 1) if the sum of its active inputs meets or exceeds its threshold. Given the graph and thresholds, determine the final state of every neuron after full propagation. ```python def propagate_neurons( n: int, edges: list[tuple[int, int]], initial_states: list[int], thresholds: list[int] ) -> list[int]: # n: number of neurons (0-indexed) # edges: list of (src, dst) directed edges # initial_states: initial 0/1 activation per neuron # thresholds: firing threshold per neuron # returns: final activation state of each neuron ... ``` ``` Example: n = 4 edges = [(0,2), (1,2), (2,3)] initial_states = [1, 1, 0, 0] thresholds = [1, 1, 2, 1] Neuron 2 receives input from 0 and 1 (both active) -> sum=2 >= threshold=2 -> fires Neuron 3 receives input from 2 (now active) -> sum=1 >= threshold=1 -> fires propagete_neurons(...) -> [1, 1, 1, 1] ``` ## Follow-ups 1. How does your solution handle neurons that have no incoming edges? 2. What if the graph had cycles - how would you detect and handle them? 3. Can a neuron's activation change more than once? Does order matter? 4. How would you extend this to support weighted edges?
## Problem Find the distance between two nodes in a tree or graph structure, likely using BFS or LCA techniques. ## Tags binary_tree, graph
## Round 1 - Coding ## Problem You are organizing a party and have collected availability windows for each guest. Find all time slots where at least `k` guests are simultaneously available. ```python def party_times( schedules: list[list[tuple[int, int]]], k: int ) -> list[tuple[int, int]]: # schedules[i] = list of (start, end) availability windows for guest i # k: minimum number of guests that must overlap # returns: merged list of time intervals where >= k guests are free ... ``` ``` Example: schedules = [ [(1, 5), (8, 10)], # Guest 0 [(2, 6)], # Guest 1 [(3, 4), (9, 11)], # Guest 2 ] k = 2 Overlaps with >= 2 guests: [2,5]: guests 0 and 1 [3,4]: guests 0, 1, and 2 [9,10]: guests 0 and 2 party_times(schedules, k) -> [(2, 5), (9, 10)] # Note: [3,4] is absorbed into [2,5] ``` ## Follow-ups 1. What is the time complexity of your approach? 2. How would you handle fractional time slots (e.g. 9:30 AM)? 3. Can you solve this using a sweep line? Walk me through the algorithm. 4. What if you need to output the exact guest lists for each slot?
## Problem Implement logic to evaluate and rank poker hands from a set of cards. ## Tags coding_other, sorting
## Round 1 - Coding / OOD ## Problem Design a survival card game where players are eliminated when their health reaches zero. Each turn a player draws a card that either deals damage to another player or heals themselves. Implement the game engine. ```python class Action: # ATTACK or HEAL action_type: str value: int class Card: def __init__(self, action: Action): ... class Player: def __init__(self, name: str, health: int = 100): ... def is_alive(self) -> bool: ... def apply_action(self, action: Action, target: 'Player' = None) -> None: ... class SurvivalGame: def __init__(self, players: list[Player]): ... def play_turn(self) -> None: ... def get_survivors(self) -> list[Player]: ... def is_over(self) -> bool: ... ``` ``` Example: p1, p2, p3 = Player("A"), Player("B"), Player("C") game = SurvivalGame([p1, p2, p3]) while not game.is_over(): game.play_turn() game.get_survivors() -> [Player("A")] # last standing ``` ## Follow-ups 1. How do you pick the target for an attack card - random, lowest health, or caller's choice? 2. How would you add special card types (e.g. shield, multi-hit)? 3. How would you log the game history for replay? 4. How does your design change to support teams?
## Problem Schedule tasks with cooldown constraints to minimize total execution time. ## Likely LeetCode equivalent LC 621 - task-scheduler ## Tags heap, greedy, sorting
## Problem Simulate a task processing system with queuing, priority, and state tracking for submitted jobs. ## Tags queue, heap, coding_other
What Scale AI Looks for in Software Engineer Interviews
Scale AI Software Engineer interviews are calibrated against the level and scope expected of the role. Across 14+ verified candidate reports on LeakCode, the consistent signals interviewers look for: clear problem decomposition before coding, explicit complexity reasoning, structured handling of edge cases, and the ability to articulate trade-offs between two reasonable approaches.
The discriminator between candidates who advance and candidates who do not is rarely the final correctness of the solution. It is the path to the solution: did you ask clarifying questions, did you state your approach before coding, did you handle edge cases without prompting, and did you communicate your reasoning throughout. Reports tagged "no hire" frequently cite a working solution with poor communication; reports tagged "strong hire" cite clear thinking even when the final solution was incomplete.
How To Use This Question Set
Real interview reports are a calibration tool, not a memorization target. Companies update their question pools every 2-4 months; memorizing exact problems risks misleading you when the interviewer uses a variant. The high-leverage use: identify the patterns that appear repeatedly in Scale AI Software Engineer reports, practice those patterns on similar (not identical) problems, and use the reports to understand the interviewer's typical follow-up depth.
Filter the questions below by round type, difficulty, and recency. Focus first on reports from the past 6-12 months; older reports may reference questions that have since rotated out of Scale AI's pool. Reports tagged with quantified difficulty (e.g., "medium-hard") are higher-signal than reports without difficulty tags.
Round-by-Round Expectations
Scale AI Software Engineer loops typically span 4-6 rounds across phone screens and on-site or virtual on-site interviews. The structure varies by company: some run 1 recruiter screen + 1 technical phone + 3-4 on-site rounds; others run 1 recruiter screen + 1 OA + 4-5 on-site rounds. The recruiter screen is logistics and culture-light; the technical phone screen is medium-difficulty coding; the on-site loop covers coding, system design (at L4+ levels), and behavioral rounds.
Each round is designed to surface a specific signal. Coding rounds: correctness, code quality, complexity reasoning, communication. System design rounds: requirements clarification, design judgment, operational thinking. Behavioral rounds: ownership scope, leadership, ambiguity tolerance, conflict navigation. Strong candidates explicitly hit each signal dimension out loud during the round; weak candidates focus only on solving the prompt.
Common Interview Mistakes At This Combination
Reports tagged "no hire" at Scale AI Software Engineer commonly cite: jumping into code without clarifying requirements, coding silently for 10+ minutes without verbalizing approach, missing edge cases (empty input, single element, very large input, overflow), and producing a working solution that the candidate cannot explain or refactor when probed. Strong candidates avoid these patterns by following a consistent template: clarify, verbalize approach, code with narration, test with examples.
Behavioral and design rounds have their own failure modes. Behavioral: stories that use "we" instead of "I" diluting individual signal, stories with no quantified outcome, defensiveness when probed about failure. Design: not asking clarifying questions, not stating requirements out loud, designing for a single server when the prompt clearly implies scale, ignoring operational concerns (deployment, monitoring, rollback). These show up in roughly half of Scale AI Software Engineer interview retrospectives on LeakCode.
See All 14 Scale AI Software Engineer Questions
Full question text, answer context, and frequency data for subscribers.
Get Access