Databricks SDE II Interview Experience: Fibonacci Trees CIDR and LLD
Question Details
Hi, I had applied for the SDE-2 role at Databricks using their career portal via LinkedIn and recruiter reached out to after couple of days and scheduled my interviews one by one. # Round - 1: Technic
Full Details
Hi, I had applied for the SDE-2 role at Databricks using their career portal via LinkedIn and recruiter reached out to after couple of days and scheduled my interviews one by one. # Round - 1: Technical Screening
Fibonacci trees are binary trees which are recursively defined as follows: * T_0 is empty * T_1 consists of a single node. * T_n consists of a root node, with T_{n-2} as its left child, and T_{n-1} as its right child. Examples: T_0: T_1: * T_2: * \ * T_3: * / \ * * \ * Now, in order to be able to identify each node within tree, we enumerated the nodes in each tree using DFS pre-order. Write a function that given nodes s and e in a Fibonacci tree of order n returns the shortest path from s to e in the form of a sequence of moves: "U" (up), "L" (left), and "R" (right).
**Example** output: fibPath(order = 3, start = 1, end = 3) == "URR" fibPath(order = 4, start = 1, end = 4) == "URL" fibPath(order = 5, start = 3, end = 7) == "UURLR"
- Successfully solved the question using pre-computing fibonacci series and then using that to find the start and end in the tree. 2. Later asked for the follow-up if we iterate the tree in in-order what difference will be in the solution. # Round - 2: DSA Given a set of rules, implement the function access_ok to see if IP address is allowed or denied. Also, print the rules index because of which it is getting allowed or denied. If none of the rules match, deny it.
vector<vector<string>> rules = { {"ALLOW", "192.168.100.0/24"}, -> 192.168.100.2 {"DENY", "192.168.0.5/30"}, {"ALLOW", "192.168.1.1/22"}, {"ALLOW", "1.0.0.0/8"}, {"ALLOW", "2.3.4.9"}, {"DENY", "8.8.8.8/1"}, {"ALLOW", "5.6.7.8"} }; access_ok(rules, ip_address) -> true/false
- Solved by generating ranges for each of the CIDR rules and then checking if the given IP lies in the range or not. Used long long to store the ranges of the IPs (can use bitset if you are comfortable with it). # Round - 3: DSA
Design Tic Tac Toe for n * m, matrix and k subsequent tokens
**Follow up** is_robot to let AI play randomly
Round - 4: LLD
You are given an interface StorageClient that allows downloading files from a remote storage system: getFileSize(uri) →
**returns** the size of the file. fetch(uri, offset, length, buffer) → fetches length bytes starting at offset into buffer. Your task is to design a class CachedFile that: - Downloads a remote file from a given uri with clients calling the CachedFile class with different start and end ranges. - Optimise the algorithm to supports repeated reads efficiently - Minimizes repeated network calls
- Solved using bucketing the ranges in small chunk size at the cache layer and storing the buckets in the disk for faster retrivel rather than network call 2. Only retrieving once for the Storage client via network call for each bucket. 3. Implemented thread pool to fetch the chunks parallely from storage client. 4. Discussed various trade offs like duplicate chunk fetching in threads, failure scenarios, etc. # Round - 5: Team Matching 1. Had multiple team matching rounds with different HMs. 2. Asked common leadership and past scenarios related questions.
Verdict Pending
About This Question
This is a reported interview question from a databricks interview for a swe role during the system design round reported in 2026.
It covers the following topics: Graph, Math, Strings, Binary Tree, Concurrency, System Design, Behavioral, Matrix .
More Databricks Interview Questions
About Databricks Interview Reports
This question was reported by a candidate who interviewed at Databricks. LeakCode aggregates interview reports from 10+ sources, including 1Point3Acres, Glassdoor, LeetCode Discuss, Blind, Reddit, Indeed, and Nowcoder. Each report is translated where necessary, deduplicated against existing entries, and tagged by company, role, round type, and reporting date.
Use this question as one calibration data point, not a memorization target. Companies typically rotate their question pools every 2-4 months; the exact wording of a 2024 question may differ from what you encounter today. The underlying pattern, difficulty level, and follow-up depth at Databricks are the higher-signal extractions to take from this report.
For broader preparation context, the Databricks interview process typically includes a recruiter screen, one or two technical phone screens, and a 4-5 round on-site loop covering coding, system design (at L4+ levels), and behavioral. Reports tagged on LeakCode show the round-by-round distribution and typical difficulty calibration. To browse questions filtered by round type and seniority, use the company hub linked above.
How To Practice This Type of Question
Solve similar problems on LeetCode under timed conditions (25-35 minutes per medium difficulty). The goal is pattern recognition: recognize the underlying technique (sliding window, two-pointer, BFS, memoized recursion, etc.) within 60-90 seconds of reading. Strong candidates verbalize their hypothesis out loud before coding, then iterate based on feedback. Weak candidates dive into implementation immediately, lose time on the wrong approach, and run out of time for follow-ups.
Companies update their question pools every 2-4 months. The exact wording of any given question may have been retired by the time you interview. Focus your prep on the pattern, not the specific problem. The patterns that appear in Databricks reports consistently are the ones worth investing in; one-off niche problems are not.
During Your Databricks Round
Apply the standard interview round template: clarify requirements (2-3 minutes), state your approach out loud and confirm direction with the interviewer (3-5 minutes), code with narration (15-25 minutes), test with concrete examples including edge cases (5 minutes), discuss optimization or trade-offs if time permits (5 minutes). This template is universally accepted across FAANG and adjacent companies; deviating from it produces weaker interviewer feedback signal.
The single most predictive failure mode in Databricks reports tagged "no hire": not asking clarifying questions. Interviewers are explicitly trained to weight this. Strong candidates ask 3-5 clarifying questions even on problems that look obvious; weak candidates dive into code immediately. The clarifying-question check is often the first signal recorded in the interviewer's written notes.