Python Interview Questions for Data Analysts
Python coding rounds evaluate whether you understand data structures, vectorization, and data wrangling with Pandas. Review these real interview coding questions tested across data analytics and business intelligence roles.
What is the fundamental difference between loc and iloc in Pandas?
loc is label-based: you select rows and columns by their index names or column labels, and slice endpoints are INCLUSIVE. iloc is integer position-based: you select rows and columns by their zero-indexed numeric positions (0, 1, 2...), and slice endpoints follow standard Python convention and are EXCLUSIVE.
# Label-based: inclusive
df.loc[0:5, ['customer_id', 'total_amount']]
# Position-based: exclusive (rows 0 to 4, columns 0 to 2)
df.iloc[0:5, 0:3]How do you detect, evaluate, and handle missing values in a Pandas DataFrame?
First evaluate null distribution with df.isna().sum(). Depending on business domain and percentage of missingness: 1) Drop rows or columns using df.dropna() if missing values are minimal (<2%) and random. 2) Impute numeric columns with median (if skewed) or mean using df["col"].fillna(df["col"].median()). 3) For time-series, use forward-fill (ffill). 4) For categorical features, impute with "Unknown" or the mode.
# Percentage of nulls per column
null_pct = df.isna().mean() * 100
# Impute skewed revenue with median
df['order_value'] = df['order_value'].fillna(df['order_value'].median())What is the operational difference between merge(), join(), and concat() in Pandas?
merge() performs SQL-style relational joins (INNER, LEFT, RIGHT, OUTER) between two DataFrames based on common key columns. join() combines DataFrames primarily on their index values. concat() stacks or appends DataFrames along an axis (axis=0 for stacking rows, axis=1 for side-by-side concatenation) without evaluating key matches.
# Relational join on common key
merged_df = pd.merge(orders_df, customers_df, on='customer_id', how='left')
# Vertical stacking of monthly log files
annual_df = pd.concat([jan_df, feb_df, mar_df], axis=0, ignore_index=True)How does groupby() with .agg() work in Pandas, and why is it superior to manual loops?
groupby() splits a DataFrame into groups based on key columns, applies aggregate functions, and combines results into a new summary DataFrame. Using .agg() allows you to calculate multiple distinct metrics on different columns simultaneously in a single pass with C-optimized speed.
summary = df.groupby('region').agg(
total_orders=('order_id', 'count'),
total_revenue=('amount', 'sum'),
avg_order_value=('amount', 'mean')
).reset_index()Practice Live Technical Assessments with SSSAM Faculty
Reading questions is the first step. Writing queries under timed screen-share tests with live mentor reviews builds genuine interview confidence.