Questions asked in real interviews at IT companies, with answers written the way you should say them in the interview room. Read the free samples below — VR IT Solutions students get the full set in the Job Ready Hub.
What is the difference between Continuous Delivery and Continuous Deployment?
Both start the same way: every change is built, tested and packaged automatically by the pipeline.
The difference is the last step:
Continuous Delivery: the release is always ready, but someone approves the push to production (for example a manual approval stage in Azure DevOps).
Continuous Deployment: if all checks pass, the change goes to production automatically with no manual step.
In interviews I usually add: most companies in regulated domains like banking use Continuous Delivery with an approval gate, while product teams with strong automated tests use Continuous Deployment.
Real-time scenario
Your deployment finished successfully in the pipeline, but the new pods are in CrashLoopBackOff. What do you check first?
I would go step by step:
kubectl describe pod <pod> — check Events for image pull errors, failed probes or OOMKilled.
kubectl logs <pod> --previous — see why the container exited on the last run.
Check ConfigMaps and Secrets — a missing environment variable or wrong connection string is the most common cause after a release.
Check liveness/readiness probes — if the app takes longer to start than the probe allows, Kubernetes keeps restarting it.
Check resource limits — if memory is too low the container is killed.
If it is a bad release, I roll back with kubectl rollout undo deployment/<name> (or redeploy the previous version from the pipeline), and then fix the issue in a lower environment.
Real-time scenario
Two engineers ran terraform apply at the same time and the state got corrupted. How do you stop this from happening again?
The fix is to move state to a remote backend with locking:
On Azure: store state in an Azure Storage account container — the azurerm backend uses blob leases for locking.
On AWS: store state in S3 with a DynamoDB table for state locking.
Then only one apply can run at a time; the second one waits or fails with a lock error.
I also run Terraform only from the pipeline (not from laptops), enable versioning on the storage so an older state can be restored, and keep separate state files per environment (dev, test, prod).
178 more questions and scenarios
Available to VR IT Solutions students in the Job Ready Hub.
Your ADF pipeline fails midway because of a transient failure in one activity. How would you design it so that the run resumes from the point of failure instead of starting from scratch?
I handle this at two levels:
Activity level: set Retry and Retry interval on the activities that call external systems, so short network or throttling errors are retried automatically.
Pipeline level: I make each step idempotent and keep a checkpoint — for example a control table or a Get Metadata check that tells the pipeline which files or tables are already loaded. An If Condition then skips the steps that finished.
ADF also has "Rerun from failed activity" in the Monitor tab, which reruns only from the failed step. In my project I combine that with the control table so even a scheduled rerun does not reload data that is already done.
Real-time scenario
You need to keep the history of changes in customer data for auditing. Which technique would you use in Databricks?
I would implement Slowly Changing Dimension Type 2 on a Delta table.
Each customer row has columns like effective_from, effective_to and is_current.
When a changed record arrives, a MERGE INTO statement closes the old row (sets effective_to and is_current = false) and inserts the new version as the current row.
Unchanged records are ignored, new customers are inserted.
Delta Lake makes this reliable because MERGE is ACID. For simple audit needs, Delta time travel and the Change Data Feed can also show what changed, but SCD Type 2 is what business users and reports usually need.
Real-time scenario
You are asked to remove nulls, fill missing values and drop duplicates from a DataFrame. How would you do it?
df_clean = (df
.dropna(subset=["id"]) # rows without a key are useless
.fillna({"amount": 0, "city": "Unknown"})
.dropDuplicates(["id"]))
I do not use dropna() on all columns blindly, because one empty optional column would delete good rows. I decide column by column: drop where the key is missing, fill defaults for numeric or text fields, and remove duplicates on the business key. I also log how many rows were dropped so data quality can be reported.
Real-time scenario
You are building a modern data warehouse in Azure. How would you explain the Medallion Architecture in your project?
Medallion architecture splits the lake into three layers:
Bronze: raw data as it arrives from the source — no changes, so we can always reprocess.
Silver: cleaned and conformed data — duplicates removed, types fixed, data joined and validated.
Gold: business-ready tables — aggregated facts and dimensions used by Power BI and reports.
In my project, ADF landed files into bronze, Databricks notebooks built silver and gold Delta tables, and Synapse or Power BI read from gold. The benefit is clear ownership of each layer and easy reprocessing when business rules change.
Real-time scenario
How do you find the 4th highest salary from an Employee table?
Using a window function:
SELECT salary
FROM (
SELECT salary, DENSE_RANK() OVER (ORDER BY salary DESC) AS rnk
FROM employee
) t
WHERE rnk = 4;
I use DENSE_RANK so that duplicate salaries do not skip a rank. If the interviewer wants the 4th row regardless of ties, ROW_NUMBER() works instead.
23 more questions and scenarios
Available to VR IT Solutions students in the Job Ready Hub.
Write a GlideRecord script to fetch all incidents created in the last 7 days.
var gr = new GlideRecord('incident');
gr.addEncodedQuery('sys_created_onONLast 7 days@javascript:gs.beginningOfLast7Days()@javascript:gs.endOfLast7Days()');
gr.query();
while (gr.next()) {
gs.info(gr.number + ' - ' + gr.short_description);
}
Tip: I usually build the filter in the list view first, right-click the breadcrumb, "Copy query", and paste it into addEncodedQuery — that way the query is exactly what the platform uses.
Name and total experience, and how much of it is on ServiceNow.
Current role and project: which modules you work on (for example ITSM — Incident, Problem, Change, Service Catalog) and which release (for example Yokohama or Zurich).
Two or three things you actually did: catalog items with variables and flows, business rules and client scripts, ACLs, notifications, an integration.
Certifications (CSA, CAD, CIS-ITSM) if you have them.
Example: "I have 3 years of experience, 2.5 years on ServiceNow ITSM. In my current project I support Incident, Change and the Service Catalog — I built 25+ catalog items with Flow Designer, wrote business rules and client scripts, and set up a REST integration with our HR system. I am a Certified System Administrator."
Do not read your resume line by line; the interviewer will pick topics from what you say, so mention what you are strong in.
Interview question · Asked at HCL
What is a Client Script? What are its types? Explain each.
A client script is JavaScript that runs in the browser on forms (and lists for onCellEdit). Types:
onLoad — runs when the form opens. Example: hide a field or show a message.
onChange — runs when a specific field value changes (and on load, with isLoading = true). Example: when Category = Hardware, show the Asset field.
onSubmit — runs when the form is saved; returning false stops the save. Example: validate that the end date is after the start date.
onCellEdit — runs when a field is edited directly in a list. Example: stop users from changing State from the list.
Catalog client scripts support onLoad, onChange and onSubmit, but not onCellEdit.
107 more questions and scenarios
Available to VR IT Solutions students in the Job Ready Hub.
SDTM (Study Data Tabulation Model) organises the collected clinical data into standard domains — DM, AE, CM, LB, VS, EX and so on. It is close to the raw CRF data, with standard variable names and controlled terminology.
ADaM (Analysis Data Model) is built from SDTM for analysis. It adds derived variables (for example AVAL, CHG, baseline flags, treatment-emergent flags) so the statistician can produce tables directly. Main datasets are ADSL (one record per subject) and BDS datasets like ADLB, ADVS.
In short: SDTM = "what was collected", ADaM = "ready for analysis". Both are required for FDA submissions along with define.xml.
Interview question
What is the difference between MERGE in a DATA step and a JOIN in PROC SQL?
DATA step MERGE needs both datasets sorted by the BY variables and matches row by row. With many-to-many matches it does not give a full cartesian product, which can give wrong results.
PROC SQL JOIN does not need sorting, supports inner, left, right and full joins, and handles many-to-many correctly (cartesian product).
In projects I use MERGE with IN= flags for simple one-to-one or one-to-many merges (for example adding DM variables to AE), and PROC SQL when keys are different names, when I need many-to-many, or summary values in the same step.
Interview question
What are TLFs and how are they validated?
TLFs are Tables, Listings and Figures — the outputs in the Clinical Study Report, for example demographics table, adverse events by system organ class, lab shift tables, Kaplan-Meier plots.
Validation is usually done by independent double programming: a second programmer writes their own program from the same specifications (SAP and mock shells), and the outputs or datasets are compared with PROC COMPARE. Any difference is discussed and fixed. We also check titles, footnotes, population counts and formatting against the mock shell.
Real-time scenario
After merging AE with DM, your AE dataset has more records than before. How do you find and fix the problem?
More records after a merge means the key is not unique in one of the datasets.
Check DM is one record per subject: PROC SORT DATA=dm NODUPKEY DUPOUT=dups; BY usubjid; — if DUPOUT has records, DM has duplicates.
Check the BY variables are correct — merging only by SUBJID instead of STUDYID and USUBJID can match subjects across studies.
Use IN= flags to see where records come from: data ae2; merge ae(in=a) dm(in=b); by usubjid; if a; run;
Compare counts before and after with PROC FREQ or a quick count.
Then I fix the source (remove the real duplicate or correct the key) rather than just dropping rows, and note it in the issue log.
Real-time scenario
The statistician asks you to derive a treatment-emergent flag (TRTEMFL) in ADAE. How would you do it?
First I read the definition in the SAP, because it can vary by study. A common rule is: the AE started on or after the first dose date, or started before but got worse after first dose.
data adae;
merge ae(in=a) adsl(keep=usubjid trtsdt trtedt);
by usubjid;
if a;
astdt = input(aestdtc, ?? yymmdd10.);
if not missing(astdt) and not missing(trtsdt) and astdt >= trtsdt then trtemfl = 'Y';
run;
I also handle partial dates according to the imputation rules in the SAP (for example missing day imputed to the first day of the month, but not before first dose), and check the flag counts with the statistician.
Real-time scenario
Pinnacle 21 shows errors in your SDTM datasets just before submission. What do you do?
Read the report and sort by severity — Errors first, then Warnings.
Group the issues: controlled terminology mismatches, missing required variables, wrong lengths or formats, ISO 8601 date problems, inconsistent USUBJIDs across domains.
Fix the root cause in the SDTM program or specification, not by editing the dataset by hand.
Re-run the programs, regenerate define.xml and run Pinnacle 21 again.
Some findings are expected because of how the data was collected — those are explained in the Reviewer's Guide (cSDRG) instead of being "fixed".
I keep a tracker of each issue, the fix, and the explanation, so the team lead can review before submission.
What is the difference between a Profile and a Permission Set?
A Profile is assigned to every user (exactly one) and sets their baseline access — object and field permissions, page layouts, login hours, IP ranges.
Permission Sets give extra access on top of the profile, and a user can have many. Permission Set Groups bundle several permission sets for a role.
Salesforce recommends keeping profiles minimal and giving access through permission sets, because it is more flexible — for example a "Discount Approver" permission set can be given to a few users without creating a new profile.
Interview question
What is bulkification and why does it matter in Apex triggers?
A trigger can receive up to 200 records at once — for example from a data load or an integration. Bulkified code handles all records in the batch together instead of one by one.
No SOQL or DML inside loops — collect Ids in a Set, query once, use a Map.
Do one DML statement on a list at the end.
This matters because of governor limits (100 SOQL queries and 150 DML statements per synchronous transaction). Code that works for one record in the UI can fail with "Too many SOQL queries" when 200 records are loaded.
Interview question
When do you use Flow and when do you use Apex?
Salesforce's guidance is "clicks before code":
Flow (record-triggered, screen, scheduled) for most business automation — field updates, creating related records, approvals, sending emails. Admins can maintain it.
Apex when the logic is complex, needs high performance on large volumes, complex integrations, or things Flow cannot do well (heavy loops, custom error handling, callouts in complex patterns).
In projects I first check if a before-save record-triggered flow can do it (fastest for same-record updates), and use Apex triggers with a handler class when the logic grows. I avoid having both a flow and a trigger updating the same field.
Real-time scenario
A data load fails with "System.LimitException: Too many SOQL queries: 101" from your trigger. How do you fix it?
This means a query is running inside a loop (or several triggers/flows fire for the same records).
Find the query inside the for loop and move it out: collect the Ids in a Set, run one query, store results in a Map<Id, SObject>.
Check for recursion — the trigger updates records that fire the same trigger again. Use a static variable or handler pattern to stop re-entry.
Check other automation on the object (flows, process builders) that add their own queries in the same transaction.
Test with 200 records in a unit test so the issue is caught before production.
Real-time scenario
Sales reps must see only Accounts in their region, but regional managers must see all accounts of their team, and the national head must see everything. How do you design it?
Set Organization-Wide Default for Account to Private, so users see only records they own.
Build the Role Hierarchy: National Head → Regional Managers → Sales Reps. With "Grant access using hierarchies", managers automatically see their reps' records and the national head sees all.
If reps in the same region need to see each other's accounts, add a criteria-based Sharing Rule (Region = South → share with role "South Sales").
Field-level access is controlled separately with profiles/permission sets.
I always test with "Login as" for one user in each role.
Real-time scenario
The business wants an AI agent that answers customers' "What is the status of my case?" questions. How would you build it with Agentforce?
Create a service agent in Agentforce Builder and add a Topic like "Case Status" with clear instructions on what it can and cannot answer.
Add Actions the agent can call — for example an autolaunched Flow (or Apex) that finds the customer's open cases by email or case number and returns status and last update.
Make sure the agent runs with a user that has only the needed access, and keep sensitive fields out of the action output.
Test conversations in the builder, including wrong case numbers and off-topic questions, then deploy to the channel (Messaging for web or Experience Cloud).
Add a handoff to a human agent when the customer is unhappy or the question is out of scope.
Explain Case Type, Stage, Process and Step in Pega.
Case Type — the business transaction being automated, for example "Loan Application".
Stages — the main phases of the case life cycle, for example Submission → Review → Approval → Fulfilment → Resolution.
Processes — inside each stage, the flows that run (one or more, in sequence or parallel).
Steps — the individual tasks inside a process: collect information (assignment), approve, send email, call an automation, wait.
In App Studio this is the Case Life Cycle view; underneath it creates flow rules, flow actions and sections.
Interview question
What is a Data Page and what are its types and scopes?
A Data Page is a clipboard page that loads data on demand from a source (report definition, connector, data transform, lookup) and caches it.
Types/modes: Read-only (most common), Editable, and Savable (can write back to the source).
Structure: Page (single record) or List.
Scope: Thread (per case/thread), Requestor (per user session), Node (shared by all users on the node — good for reference data like country lists).
Refresh strategy (reload if older than X, or a when condition) controls when the cache is reloaded. Using data pages avoids loading the same data again and again in activities.
Interview question
How does Pega rule resolution work?
When Pega needs a rule, it picks the best one in this order:
Checks the rule cache first.
Filters by the operator's ruleset stack (application rulesets and versions) and the rule's availability (blocked, withdrawn, final, available).
Uses class inheritance — starts in the case class and moves up through pattern inheritance and then direct inheritance until it finds the rule.
Applies circumstancing (property or date circumstances) to pick the special version if it applies.
Checks access (privileges) and returns the best match.
Knowing this helps to debug "why is my rule not picked up" problems.
Real-time scenario
A request should go to the reporting manager for approval only if the amount is above 50,000; otherwise it should be auto-approved. How do you implement it?
Add an Approval step in the stage and set it to "Approve/Reject" with routing to "Reporting manager".
Put a condition on that step (a When rule like Amount > 50000), so the step runs only when the condition is true; otherwise the case moves on.
For the auto-approved path, set a status like "Resolved-Approved" or continue to fulfilment.
If the threshold may change, store 50,000 in a data page or decision table instead of hard-coding it, so business users can change it.
I would test both paths with the Tracer and check the audit history.
Real-time scenario
Cases not resolved within 2 days must be escalated to the team lead. How do you configure this?
Create a Service Level Agreement rule and apply it to the assignment or the case:
Goal — for example 1 day: sends a reminder to the assignee.
Deadline — 2 days: escalation action, for example raise urgency and notify the team lead.
Passed deadline — repeat or transfer the assignment to the team lead's workbasket.
Urgency increments push the case up the worklist. In Pega 8 the SLA events are processed in the background by the ServiceLevelEvents queue processor, so I make sure it is running in each environment.
Real-time scenario
Users complain a Pega screen takes 10 seconds to load. How do you investigate?
Reproduce with Tracer and the Performance tool (PAL) to see which steps and DB calls take the time.
Check the DB Trace for slow queries — often a report definition without filters or indexes.
Look for data pages loading too often (wrong scope or refresh strategy) or large lists loaded on screen load.
Check sections for "refresh on change" causing repeated server round trips.
Check PDC (Predictive Diagnostic Cloud) alerts if available.
Then fix the cause — add filters/indexes, change data page scope, load lists lazily — and measure again.
map loops over the array, the object constructor builds the new shape, and reduce sums the line totals. I also use filter, groupBy and default values (order.customer.name default "") in real transformations.
Real-time scenario
A backend API your Mule app calls fails intermittently. How do you make the integration reliable?
Set sensible connection and response timeouts on the HTTP Request.
Wrap the call in Until Successful with a limited number of retries and a delay, only for errors that are temporary (timeouts, 503).
Use On Error Propagate or On Error Continue deliberately: return a clear error to the caller, or continue with a fallback.
For non-urgent requests, put the message on a queue (Anypoint MQ/VM/JMS) and process it asynchronously, so a failure does not lose data.
Log with a correlation ID and set up alerts in Anypoint Monitoring.
Real-time scenario
You need to process a CSV file with 1 million records into a database. How would you design it in Mule?
Read the file with streaming so the whole file is not loaded into memory.
Use a Batch Job: the input phase splits records, batch steps validate and transform each record, and a Batch Aggregator inserts records in bulk (for example 500 at a time) using a bulk insert.
Use the batch "accept policy" to send failed records to an error step and write them to an error file or table.
In the On Complete phase, log totals (processed, successful, failed) and send a summary.
Tune block size and max concurrency based on the database capacity.
Real-time scenario
How do you secure a Mule API that external partners will call?
Apply policies in API Manager: Client ID enforcement for each partner app, OAuth 2.0 token enforcement if needed, rate limiting/SLA tiers, and IP allowlist.
Use HTTPS with TLS on the listener.
Keep credentials in secure properties (encrypted) and never in plain config files.
Validate the request against the RAML/OAS spec with APIkit so bad payloads are rejected early.
Log access without sensitive data, and monitor usage per client in Anypoint.
What is a Supervisory Organization in Workday and how is it different from a Cost Center?
A Supervisory Organization is the backbone of Workday HCM — it groups workers who report to the same manager, and drives business process routing and security (who approves, who sees what).
A Cost Center is a financial organization used for costing and reporting; it does not decide reporting lines.
A worker belongs to one supervisory org but can be assigned to cost centers, companies, regions and other custom organizations for reporting and accounting.
Interview question
What is the difference between Position Management and Job Management?
They are the two main staffing models, set per supervisory organization:
Position Management — every worker fills a pre-defined position; you must create the position before hiring. Gives tight headcount control and budgeting.
Job Management — you define job profiles and a headcount limit, and hire against the org without creating positions one by one. More flexible.
There is also Headcount Management, which controls headcount numbers without individual positions.
Interview question
What is a calculated field? Give an example.
A calculated field creates a new value from existing fields without changing the data, using functions like Arithmetic Calculation, Date Difference, True/False condition, Lookup Related Value or Concatenate Text.
Example: "Years of Service" = Date Difference between Hire Date and Today, in years. It can then be used in reports, conditions in business processes, and integrations. Good practice is to name and describe calculated fields clearly, because many reports may depend on them.
Real-time scenario
The Hire business process should need HR Partner approval only for hires in Germany. How do you configure it?
Edit the Hire business process definition for the relevant org (or the default definition).
Add or find the Approval step for HR Partner.
Add a condition (entry condition) on that step: Location → Country is Germany. Create or reuse a condition rule / calculated field if needed.
Save and test with one hire in Germany and one elsewhere, checking the process history to confirm the step is skipped when it should be.
I document the change and test in a sandbox before moving it to production.
Real-time scenario
You need to load address changes for 500 workers. What is the easiest way?
Use an EIB (Enterprise Interface Builder):
Create an inbound EIB using the web service "Change Home Contact Information".
Generate the spreadsheet template, fill one row per worker with Employee ID and the new address fields.
Upload the spreadsheet and launch the EIB, first in a sandbox.
Check the results — the EIB shows errors per row (for example missing postal code) so you can fix and reload only those.
For regular changes from another system, I would build an integration instead of a one-time EIB.
Real-time scenario
A third-party payroll provider needs only the changes in worker data every pay period. Which integration approach would you use?
I would use a Workday payroll connector such as PECI (Payroll Effective Change Interface) or PICOF, which are designed for third-party payroll and send effective-dated changes only.
Configure the connector with the payroll fields needed (personal data, compensation, earnings and deductions).
Transform the output to the vendor's format with an XSLT or Document Transformation if required.
Schedule it per pay period and deliver by SFTP.
For non-payroll vendors needing changes, Core Connector: Worker with change detection is the usual choice.
Storage — data is stored in compressed columnar micro-partitions in cloud storage (S3, Azure Blob or GCS).
Compute — virtual warehouses (clusters) run queries. Each warehouse is independent, so ETL and BI can use different warehouses without slowing each other.
Cloud services — handles metadata, query optimisation, security and transactions.
Because storage and compute are separate, you can scale compute up or out, and pay for storage and compute separately.
Interview question
What are Time Travel and Fail-safe?
Time Travel lets you query, clone or restore data as it was at an earlier point — for example SELECT * FROM orders AT(OFFSET => -3600) or UNDROP TABLE orders. The retention is 1 day by default (up to 90 days on Enterprise edition).
Fail-safe is a further 7-day period after Time Travel ends, during which Snowflake support can recover data. Users cannot query it directly.
Both add storage cost, so for temporary or staging tables we often use transient tables with low retention.
Interview question
What is zero-copy cloning?
CREATE TABLE orders_dev CLONE orders; creates a copy instantly without copying the data — both point to the same micro-partitions. Only changes made after the clone use new storage.
You can clone tables, schemas and whole databases, and combine with Time Travel (clone as of a past time). We use it to create dev/test copies of production data in seconds.
Real-time scenario
Files arrive continuously in an S3 bucket and must be loaded into Snowflake within minutes. How do you design it?
Use Snowpipe with auto-ingest:
Create a storage integration and an external stage pointing to the S3 path.
Create a file format (CSV, JSON or Parquet) and a pipe: CREATE PIPE ... AUTO_INGEST = TRUE AS COPY INTO raw_orders FROM @s3_stage;
Configure S3 event notifications (SQS) so Snowpipe is triggered when a new file lands.
Monitor with COPY_HISTORY and handle bad records with ON_ERROR options.
Then a Stream and Task (or dynamic tables) move data from raw to clean tables.
Real-time scenario
Queries are slow and credit usage is high. What do you check?
Query Profile — look for large table scans, spilling to disk, or exploding joins.
Warehouse size and settings — right-size it, set AUTO_SUSPEND to 60 seconds and AUTO_RESUME on; use multi-cluster only for concurrency.
Pruning — make filters use columns that prune micro-partitions; add a clustering key on very large tables filtered by date.
Avoid SELECT * and repeated heavy transformations — materialise results or use the result cache.
Separate warehouses for ETL and BI, and set resource monitors to alert on credit usage.
Real-time scenario
You need to process only new and changed rows from a table every 15 minutes. How?
Use a Stream and a Task:
CREATE STREAM orders_stream ON TABLE raw_orders;
CREATE TASK merge_orders
WAREHOUSE = etl_wh
SCHEDULE = '15 MINUTE'
WHEN SYSTEM$STREAM_HAS_DATA('ORDERS_STREAM')
AS
MERGE INTO orders t USING orders_stream s ON t.id = s.id
WHEN MATCHED THEN UPDATE SET t.status = s.status
WHEN NOT MATCHED THEN INSERT (id, status) VALUES (s.id, s.status);
ALTER TASK merge_orders RESUME;
The stream tracks changes since it was last consumed, and the WHEN condition avoids starting the warehouse when there is no new data.
RAG (Retrieval-Augmented Generation) means: before the LLM answers, we search our own documents for the most relevant pieces and put them in the prompt, so the model answers from that context.
Why: LLMs do not know your company's private or latest data, and they can make things up. RAG gives up-to-date, source-based answers without retraining the model, and lets you show citations. Typical pieces: document chunking, embeddings, a vector database, retrieval (often hybrid search + reranking), and a prompt that tells the model to answer only from the context.
Interview question
What do temperature and top-p control?
Both control how random the model's output is.
Temperature — low (0–0.3) gives focused, repeatable answers; high (0.8+) gives more varied, creative text.
Top-p (nucleus sampling) — the model picks only from the smallest set of tokens whose total probability is p (for example 0.9).
For data extraction, classification and code I use low temperature; for brainstorming or marketing text, higher. Usually you tune one of them, not both.
Interview question
When do you use prompt engineering, RAG or fine-tuning?
Prompt engineering — first choice: clear instructions, examples and output format. Cheap and fast.
RAG — when the model needs your own or changing knowledge (policies, product docs, tickets).
Fine-tuning — when you need a consistent style or format, or a specific task done many times where prompting is not enough, and you have good training examples.
In practice most business apps use good prompts + RAG; fine-tuning is used less often because it costs more and must be redone when data changes.
Real-time scenario
Your company chatbot gives confident but wrong answers. How do you reduce hallucinations?
Ground it with RAG — retrieve relevant documents and instruct the model to answer only from them, and to say "I don't know" when the answer is not in the context.
Improve retrieval — better chunking, hybrid search, reranking — because many wrong answers come from wrong context.
Show citations so users can check the source.
Lower temperature and use structured output for factual tasks.
Build an evaluation set of real questions with correct answers and test every change against it.
Real-time scenario
You must build a Q&A assistant over 10,000 internal PDFs. How would you design it?
Ingestion: extract text (with OCR for scanned files), clean it, split into chunks of a few hundred tokens with some overlap, and keep metadata (document name, page, department, date).
Embeddings: create embeddings for each chunk and store them in a vector database (or a search service with vector + keyword search).
Retrieval: hybrid search, filter by metadata and user permissions, take the top results and rerank.
Generation: a prompt that includes the retrieved chunks, asks for an answer with citations, and says when information is missing.
Operations: re-index when documents change, log questions and feedback, and evaluate answer quality regularly.
Real-time scenario
Your Gen AI app works well but costs and response times are too high. What do you do?
Use a smaller or cheaper model for simple steps (classification, routing) and the large model only where needed.
Azure OpenAI Service — GPT and embedding models with Azure security and networking.
Azure AI Language — sentiment, key phrases, entity recognition, summarisation, question answering, conversational language understanding.
Azure AI Vision — image analysis, OCR, face (limited access).
Azure AI Speech — speech to text, text to speech, translation.
Azure AI Document Intelligence (earlier Form Recognizer) — extract fields and tables from documents.
Azure AI Search — keyword, vector and hybrid search, often used for RAG.
They are managed through Azure AI Foundry and secured with keys or Microsoft Entra ID.
Interview question
What is the difference between Azure OpenAI and using the OpenAI API directly?
The models are similar, but Azure OpenAI gives enterprise features:
Runs in your Azure subscription and region, with private endpoints and VNet support.
Microsoft Entra ID authentication and managed identities instead of only API keys.
Built-in content filtering and Azure compliance commitments; customer data is not used to train the models.
Integration with other Azure services like AI Search, Key Vault and Monitor.
Companies already on Azure usually choose Azure OpenAI for security and governance reasons.
Interview question
What is the difference between prebuilt and custom models in Document Intelligence?
Prebuilt models work out of the box for common documents — invoices, receipts, ID documents, business cards, tax forms, and a general layout/read model.
Custom models are trained on your own document type (for example a specific insurance claim form) with a few labelled samples: template models for fixed layouts, neural models for varied layouts.
I start with a prebuilt model if one fits, and train a custom model only when fields are missing or the format is unique.
Real-time scenario
Your client receives thousands of scanned invoices and wants the fields extracted automatically. How do you build it?
Use the Document Intelligence prebuilt invoice model to extract vendor, invoice number, dates, totals and line items.
Trigger it automatically — for example a Logic App or Azure Function when a file lands in Blob Storage.
Check the confidence score of each field; low-confidence results go to a human review queue.
Save results to a database or ERP via API, and keep the original file link.
If the vendor formats are unusual, train a custom model and compose it with the prebuilt one.
Real-time scenario
Build an internal chatbot that answers from company documents on Azure. What is your design?
Store documents in Blob Storage or SharePoint and index them in Azure AI Search with vector + keyword (hybrid) search and semantic ranking.
Use Azure OpenAI for embeddings and for the chat model; the app retrieves the top chunks and sends them in the prompt ("on your data" pattern).
Secure everything with managed identities, private endpoints and Entra ID login, and apply document-level permissions so users see only what they are allowed to.
Add content safety filters, logging in Azure Monitor and an evaluation set to check answer quality.
Real-time scenario
How do you secure keys and access for Azure AI services in production?
Prefer Microsoft Entra ID with managed identities over API keys — the app's identity gets a role like "Cognitive Services OpenAI User".
If keys are needed, store them in Azure Key Vault and rotate them; never in code or config files.
Disable public network access and use private endpoints inside a VNet.
Apply least-privilege RBAC and monitor usage and costs with Azure Monitor and alerts.
What are the Claude model families and when do you use each?
Anthropic offers models in tiers:
Opus — the most capable, for complex reasoning, long multi-step work and hard coding tasks.
Sonnet — the balanced choice for most professional work: writing, analysis, coding and agents, at good speed and cost.
Haiku — the fastest and cheapest, for simple high-volume tasks like classification, extraction and routing.
In projects I use Sonnet by default, Haiku for simple steps in a pipeline, and Opus only where the extra quality is worth the cost. Model versions change often, so I check the current list in Anthropic's documentation.
Interview question
What makes a good prompt for Claude?
Give context: who the audience is, what the output is for.
State the task clearly and what "good" looks like.
Give examples of the input and the expected output when format matters.
Separate instructions from data — for example put documents inside XML tags like <document>…</document>.
Ask for a specific output format (table, JSON, bullet list) and length.
For complex tasks, break it into steps or let the model think before answering.
Then test the prompt on several real inputs and improve it, instead of judging from one try.
Interview question
What is MCP (Model Context Protocol)?
MCP is an open standard from Anthropic for connecting AI applications to tools and data. An MCP server exposes tools (actions), resources (data) and prompts; an MCP client such as Claude Desktop, Claude Code or an agent built with the SDK can then use them.
Example: an MCP server for Jira lets Claude search issues and create tickets. The benefit is that one connector works with any MCP-compatible client, instead of building a custom integration for each app.
Real-time scenario
Your team spends 3 hours every Monday preparing a weekly status report from Excel and emails. How would you automate it with Claude?
First map the workflow: which files and emails, what the report must contain, who receives it.
Create a Claude Project with the report template, an example of a good past report, and clear instructions.
Each week, upload the Excel and paste or connect the email updates; ask Claude to summarise by project, highlight risks and fill the template.
Save the prompt as a reusable template (or a Skill) so anyone can run it.
Next step: schedule it — a script or a no-code tool (n8n/Zapier/Make) collects the files and calls the Claude API, then a person reviews before sending.
Measure the time saved and the edits needed, and improve the prompt.
Real-time scenario
You are calling the Claude API from Python and responses are slow and costly. How do you optimise?
Use the right model — Haiku for simple steps, Sonnet for the main task.
Shorten the prompt: remove repeated instructions and send only the needed context.
Use prompt caching for long system prompts or documents that repeat across calls.
Set max_tokens to a sensible limit and ask for concise, structured output.
Stream the response so users see text immediately.
Use the Message Batches API for large non-urgent jobs, which is cheaper.
Log token usage per request to see where cost comes from.
Real-time scenario
Build an AI agent that helps the DevOps team troubleshoot failed pipelines. How would you design it?
Goal: given a failed pipeline run, find the likely cause and suggest a fix.
Tools: read pipeline logs (CI API), read recent commits, query Kubernetes events, search the runbook/knowledge base — exposed as MCP servers or SDK tools.
Loop: the agent reads the error, decides which tool to call, gathers evidence, and writes a summary with the root cause and suggested fix.
Guardrails: read-only access by default; any action like re-running or rolling back needs human approval.
Testing: use real past failures as test cases and check whether the agent finds the right cause.
Built with the Claude Agent SDK, it can post its summary to Slack or Teams.
What is the difference between a list, a tuple, a set and a dictionary in Python?
List — ordered, changeable, allows duplicates: [1, 2, 2].
Tuple — ordered, cannot be changed after creation: (1, 2). Good for fixed records and as dictionary keys.
Set — unordered, unique values only: {1, 2}. Fast membership checks and removing duplicates.
Dictionary — key-value pairs: {"name": "Ravi", "age": 30}. Fast lookups by key.
In AI projects I use lists for batches of data, dicts for JSON-like records and configs, and sets to remove duplicates.
Interview question
What is overfitting and how do you prevent it?
Overfitting means the model learns the training data too well, including noise, so it scores high on training data but poorly on new data.
Prevention:
Use a separate validation/test set or cross-validation to detect it.
Get more data or use data augmentation.
Simplify the model or use regularisation (L1/L2, dropout for neural networks).
Early stopping when validation loss stops improving.
Feature selection to remove noisy features.
Interview question
What is LangChain (or LangGraph) used for?
LangChain is a Python framework for building LLM applications. It gives ready building blocks: prompt templates, connectors to LLMs, document loaders, text splitters, vector store integrations, retrievers and tools — so building a RAG app or a chatbot needs less code.
LangGraph, from the same team, is for agent workflows as a graph of steps with state, loops and human-in-the-loop checkpoints. I use LangChain for RAG pipelines and LangGraph when the agent needs multiple steps and decisions.
Real-time scenario
You receive a CSV with missing values, duplicates and wrong data types before training a model. How do you clean it in pandas?
I first check df.info() and df.isna().sum() to understand the problems, and I log how many rows were dropped or filled so the result can be explained.
Real-time scenario
Build a simple RAG chatbot in Python over company policy PDFs. What steps would you follow?
Load PDFs and extract text (pypdf or a document loader).
Split into chunks (for example 500–800 tokens with overlap).
Create embeddings and store them in a vector database (FAISS, Chroma or a managed one).
On each question: embed the question, retrieve the top chunks, build a prompt with those chunks and instructions to answer only from them with the source name.
Call the LLM and return the answer with citations.
Expose it with FastAPI and a simple UI, and test with a list of real employee questions.
Real-time scenario
Your AI agent sometimes loops calling the same tool again and again. How do you fix it?
Set a maximum number of steps/tool calls and stop with a clear message when reached.
Improve tool descriptions so the model knows exactly when to use each tool and what it returns.
Return useful results or clear errors from tools — an empty or vague result often makes the agent retry.
Keep state of what was already tried and include it in the next step.
For critical flows, use a graph (LangGraph) with defined steps instead of a fully free loop.
Boto3 is the AWS SDK for Python. It lets Python code create and manage AWS resources.
import boto3
s3 = boto3.client("s3")
s3.upload_file("report.csv", "my-bucket", "reports/report.csv")
for obj in s3.list_objects_v2(Bucket="my-bucket", Prefix="reports/").get("Contents", []):
print(obj["Key"])
Credentials come from an IAM role (on EC2/Lambda) or the AWS CLI profile — never hard-coded in the script.
Interview question
What is AWS Lambda and what are its limits you should know?
Lambda runs your code without managing servers, triggered by events (API Gateway, S3, SQS, EventBridge schedules). You pay per request and per execution time.
Key limits: maximum 15 minutes per run, memory from 128 MB to 10 GB (CPU grows with memory), deployment package size limits (use layers or container images for big libraries), and cold starts on the first call. For long jobs I use Step Functions, ECS/Fargate or AWS Batch instead.
Interview question
What is the difference between a process and a thread in Python, and what is the GIL?
A thread runs inside a process and shares its memory; a process has its own memory.
The GIL (Global Interpreter Lock) in CPython allows only one thread to execute Python bytecode at a time.
So threads are good for I/O-bound work (API calls, file or S3 operations) because they wait most of the time; for CPU-heavy work I use multiprocessing (separate processes) or libraries that release the GIL (NumPy). For many AWS API calls, threads or asyncio speed things up a lot.
Real-time scenario
When a file is uploaded to S3, it must be validated and loaded into DynamoDB. How do you build it?
Configure an S3 event notification (ObjectCreated) to trigger a Lambda function.
The Lambda reads the bucket and key from the event, downloads the file with boto3, validates each row.
Valid rows are written with a DynamoDB batch_writer; invalid rows go to an "errors/" prefix or an SQS dead-letter queue.
Give the Lambda an IAM role with only s3:GetObject on that bucket and dynamodb:BatchWriteItem on that table.
Log to CloudWatch and alarm on errors. For very large files, use Step Functions or Glue instead of one Lambda.
Real-time scenario
Write a Python script that stops all EC2 instances tagged Environment=dev every night.
import boto3
ec2 = boto3.client("ec2")
resp = ec2.describe_instances(Filters=[
{"Name": "tag:Environment", "Values": ["dev"]},
{"Name": "instance-state-name", "Values": ["running"]},
])
ids = [i["InstanceId"] for r in resp["Reservations"] for i in r["Instances"]]
if ids:
ec2.stop_instances(InstanceIds=ids)
print("Stopped:", ids)
I deploy it as a Lambda and schedule it with an EventBridge rule (for example 9 PM IST). The role needs only ec2:DescribeInstances and ec2:StopInstances. This saves a lot of cost on dev environments.
Real-time scenario
Your Lambda works locally but fails in AWS with "AccessDenied". What do you check?
The Lambda execution role — does it have the exact permission (for example s3:GetObject) on the right resource ARN?
Resource policies — S3 bucket policy, KMS key policy (encrypted objects need kms:Decrypt), DynamoDB or SQS policies.
Region and account — is the code pointing to a resource in another region or account?
SCPs or permission boundaries in AWS Organizations blocking the action.
CloudTrail shows the denied call and which principal made it — that is the fastest way to find the missing permission.
Reinforcement — an agent learns by trial and error with rewards. Examples: game playing, robotics, recommendation tuning.
Interview question
When is accuracy a bad metric? What do you use instead?
Accuracy is misleading when classes are imbalanced. If only 1% of transactions are fraud, a model that always says "not fraud" is 99% accurate but useless.
Better metrics: precision (of predicted frauds, how many are real), recall (of real frauds, how many we caught), F1-score, ROC-AUC or PR-AUC, and the confusion matrix. The choice depends on cost — for fraud or disease detection we usually care more about recall.
Interview question
What is the difference between bagging and boosting?
Both combine many models (usually decision trees):
Bagging (Random Forest) — trains trees in parallel on random samples and averages them. Reduces variance (overfitting).
Boosting (XGBoost, LightGBM, AdaBoost) — trains trees one after another, each fixing the errors of the previous ones. Reduces bias and often gives top accuracy on tabular data, but needs careful tuning to avoid overfitting.
Real-time scenario
You are building a fraud model and only 0.5% of records are fraud. How do you handle it?
Use the right metrics — precision, recall, PR-AUC — not accuracy.
Use stratified train/test splits so both have fraud cases.
Handle imbalance: class weights (class_weight="balanced" or scale_pos_weight in XGBoost), undersampling the majority, or SMOTE on the training set only.
Tune the decision threshold based on the business cost of false positives vs missed fraud.
Validate on a time-based split if fraud patterns change over time.
Real-time scenario
Your model performed well in testing but accuracy dropped after 3 months in production. Why, and what do you do?
Most likely data drift or concept drift — the input data or the relationship between inputs and outcome has changed (new customer behaviour, new products, seasonality).
Monitor input feature distributions and prediction distributions against training data.
Track actual outcomes when they arrive and compute live metrics.
Retrain on recent data, possibly on a schedule, and compare with the current model before replacing it.
Version data and models (MLflow) so you can roll back.
Real-time scenario
You have a date column and a city column. How do you prepare them for a model?
Date: extract useful parts — day of week, month, weekend flag, days since an event, holiday flag. Do not feed the raw date string.
City: for few categories use one-hot encoding; for many categories use target encoding (with cross-validation to avoid leakage) or frequency encoding; tree models like LightGBM/CatBoost can handle categoricals directly.
Fit encoders on training data only and apply the same transformation to test and production data — a scikit-learn Pipeline helps.
The p-value is the probability of seeing results at least as extreme as ours if the null hypothesis were true.
If p is smaller than the chosen significance level (often 0.05), we reject the null hypothesis. Example: an A/B test of a new checkout page gives p = 0.01, so the difference in conversion is unlikely to be random.
A p-value does not tell the size or business importance of the effect, so I also report the effect size and confidence interval.
Interview question
What is the difference between correlation and causation?
Correlation means two variables move together; causation means one actually causes the other.
Example: ice cream sales and drowning cases are correlated because both rise in summer — ice cream does not cause drowning.
To claim causation we need controlled experiments (A/B tests) or careful causal methods that handle confounding variables. In reports I say "associated with" unless there is an experiment.
Interview question
Explain the steps of a data science project.
Understand the business problem and define success (metric and target).
Collect data and understand it (EDA — distributions, missing values, outliers).
Clean and prepare data, engineer features.
Build baseline and then better models; validate properly.
Interpret results and check fairness and errors.
Deploy (API, batch scores or dashboard) and monitor.
Communicate results to business users in simple terms and measure the impact.
Real-time scenario
Marketing says the new email design increased clicks from 4% to 4.5%. How do you check if it is real?
Check the test was set up properly — random assignment, same time period, enough sample size.
Run a two-proportion z-test (or chi-square) on clicks vs sends for both groups and look at the p-value and confidence interval of the difference.
Check practical significance — is a 0.5 point lift worth the change?
Look for problems: different audiences, other campaigns running, early stopping.
Then report: "the lift is X% with a 95% confidence interval of A–B" in simple words.
Real-time scenario
You get a new sales dataset and the business wants insights by tomorrow. How do you approach it?
Clarify the questions first — revenue trend? best products? regions?
Quick data checks: shape, data types, missing values, duplicates, date range, obvious errors (negative quantities).
Summary statistics and simple charts: sales over time, top products and regions, distributions.
Look for outliers and seasonality, and segment (new vs returning customers).
Present 3–5 key findings with charts and clear next steps, and mention data limitations.
Real-time scenario
Predict which customers will churn next month. Walk through your approach.
Define churn clearly (for example no purchase or cancelled subscription in the next 30 days).
Build features from history: recency, frequency, monetary value, complaints, usage trend, plan changes.
Create a time-based training set — features as of a cut-off date, label from the following month.
Train models (logistic regression baseline, then gradient boosting), evaluate with recall/precision and AUC.
Explain drivers with feature importance or SHAP, and give the business a ranked list of high-risk customers for retention action.
The Project is the main unit: billing, APIs and IAM are managed per project.
Folders group projects (for example by department or environment).
IAM policies set at a higher level are inherited by everything below.
I usually keep separate projects for dev, test and prod, and apply organisation policies (for example restrict regions) at the top.
Interview question
What is the difference between a user account and a service account in GCP?
A user account is a person (Google account or Cloud Identity).
A service account is an identity for applications and VMs — for example a Compute Engine VM or Cloud Run service uses a service account to call Cloud Storage or BigQuery.
Best practice: give service accounts only the roles they need, avoid downloading service account keys (use attached service accounts or Workload Identity Federation instead), and use predefined or custom roles rather than broad basic roles like Editor.
Interview question
When do you choose Compute Engine, GKE, Cloud Run or App Engine?
Compute Engine — full VMs, when you need OS control or lift-and-shift.
GKE — managed Kubernetes for many containerised microservices and teams already using Kubernetes.
Cloud Run — serverless containers that scale to zero; great for APIs and web apps without managing clusters.
App Engine — older PaaS for web apps in supported runtimes.
For new containerised apps I usually start with Cloud Run and move to GKE only when more control is needed.
Real-time scenario
A VM in a private subnet with no external IP needs to download updates from the internet. How do you set it up?
Use Cloud NAT:
Create a Cloud Router in the VPC region.
Create a Cloud NAT gateway on that router for the private subnet.
The VM can then reach the internet for outbound connections (updates, external APIs) while staying unreachable from the internet. For Google APIs only (Cloud Storage, BigQuery), enable Private Google Access on the subnet instead.
Real-time scenario
The GCP bill jumped this month. How do you find out why and control it?
Look at Billing reports grouped by project, service and SKU to find what increased.
Export billing to BigQuery for detailed analysis (labels, resources).
Common causes: VMs left running, oversized machines, BigQuery queries scanning full tables, network egress, logging volume.
Fixes: rightsizing recommendations, committed use discounts, scheduling dev VMs to stop, partitioned BigQuery tables, log exclusions.
Set budgets and alerts per project so the team is notified early next time.
Real-time scenario
A developer accidentally committed a service account key to GitHub. What do you do?
Immediately disable and delete that key in IAM.
Check Cloud Audit Logs for any activity using that key and look for unknown resources.
Rotate any other secrets that might have been exposed.
Remove the key from the repository history.
Prevent it next time: use attached service accounts or Workload Identity Federation instead of keys, store secrets in Secret Manager, enable secret scanning, and use the organisation policy that blocks service account key creation.
What is the difference between partitioning and clustering in BigQuery?
Partitioning splits a table into segments, usually by a date/timestamp column or ingestion time (or integer range). Queries filtering on the partition column scan only the needed partitions — less cost and faster.
Clustering sorts data within partitions by up to four columns (for example customer_id, region). Filters on those columns skip blocks of data.
Common design: partition by event date and cluster by the most filtered columns. I also set "require partition filter" on large tables so nobody scans everything by mistake.
Interview question
What is Cloud Composer and when do you use it?
Cloud Composer is Google's managed Apache Airflow. You write pipelines as DAGs in Python, with tasks like loading files to BigQuery, running SQL, triggering Dataflow or Vertex AI jobs.
Use it when you have many dependent steps, schedules, retries and monitoring needs across services. For a single scheduled query, BigQuery scheduled queries are simpler; for event-driven small steps, Cloud Functions/Workflows can be enough.
Interview question
What does Vertex AI provide for ML projects?
Vertex AI is GCP's ML platform:
Workbench notebooks for development.
Training — custom training jobs with prebuilt or custom containers, and AutoML.
Model Registry and Endpoints for deployment (online) and batch prediction.
Pipelines (Kubeflow-based) for repeatable ML workflows.
Feature Store, model monitoring for drift, and Generative AI with Gemini models, embeddings and vector search.
Real-time scenario
An analyst's daily query scans 5 TB and is expensive. How do you reduce the cost?
Check the query plan and bytes processed.
Select only needed columns — never SELECT * in BigQuery, cost depends on columns scanned.
Make sure the table is partitioned and the query filters on the partition column; add clustering on common filters.
Create a smaller aggregated table or materialised view that the dashboard reads instead of the raw table.
Use BI Engine or cached results for dashboards, and set custom quotas to prevent runaway costs.
Real-time scenario
Daily files land in Cloud Storage and must be loaded into BigQuery without duplicates, even if a file is re-sent. How?
Load each file first into a staging table (bq load / load job from Composer).
MERGE from staging into the final table on the business key:
MERGE dataset.orders T
USING dataset.orders_stg S
ON T.order_id = S.order_id
WHEN MATCHED THEN UPDATE SET status = S.status, updated_at = S.updated_at
WHEN NOT MATCHED THEN INSERT ROW;
Keep a file-tracking table (file name, load time, row count) to skip files already processed.
Orchestrate with Cloud Composer with retries and data quality checks.
Real-time scenario
Your Vertex AI model's predictions are getting worse. How do you set up monitoring and retraining?
Enable Vertex AI Model Monitoring on the endpoint to detect training-serving skew and prediction drift on key features, with alert thresholds.
Log predictions and, when actual outcomes arrive, compute real accuracy in BigQuery.
Build a Vertex AI Pipeline that retrains on recent data, evaluates against the current model, and registers/deploys only if it is better.
Trigger it on a schedule from Cloud Composer or when drift alerts fire.
Keep model versions in Model Registry so you can roll back quickly.
NameNode — the master that keeps metadata: file names, blocks and which DataNodes hold them.
DataNodes — store the actual data blocks (default 128 MB) and send heartbeats to the NameNode.
Replication — each block is stored on 3 DataNodes by default, so a node failure does not lose data.
For high availability there is an active and a standby NameNode with JournalNodes.
Interview question
What does YARN do?
YARN (Yet Another Resource Negotiator) manages cluster resources and schedules jobs.
ResourceManager — allocates CPU/memory across the cluster.
NodeManager — runs on each node and launches containers.
ApplicationMaster — one per job, requests containers and tracks the job.
It lets different engines like MapReduce, Spark and Hive (on Tez) share the same cluster.
Interview question
What is the difference between managed and external tables in Hive?
Managed (internal) table — Hive owns the data; DROP TABLE deletes both the metadata and the data files.
External table — Hive only stores metadata pointing to an HDFS/S3 location; DROP TABLE removes only the metadata and the files remain.
I use external tables for raw data shared with other tools, and managed tables for intermediate data Hive fully controls. Partitioning (for example by date) speeds up queries on both.
Real-time scenario
A Hive query on a large table takes hours. How do you speed it up?
Partition the table by commonly filtered columns (date) and make queries filter on them.
Use columnar formats like ORC or Parquet with compression.
Bucket tables used in joins, and use map-side joins for small tables.
Run on Tez or Spark instead of MapReduce, and enable vectorisation and cost-based optimisation with table statistics.
Avoid too many small files — compact them.
Real-time scenario
HDFS has millions of small files and the NameNode is running out of memory. What do you do?
Each file and block uses NameNode memory, so many small files are a problem.
Combine small files into larger ones — Hive/Spark compaction jobs, HAR archives, or SequenceFiles.
Fix the source — make ingestion write bigger files (batch writes, fewer reducers outputs, coalesce in Spark).
Use columnar formats (ORC/Parquet) with sensible file sizes close to the block size.
Monitor file counts per directory to catch the problem early.
Real-time scenario
Your company wants to move from an on-prem Hadoop cluster to the cloud. What would you plan?
Inventory data, Hive tables, jobs (Hive, Spark, Sqoop, Oozie) and dependencies.
Choose the target: managed Spark/Hadoop (EMR, Dataproc, HDInsight) or a lakehouse (Databricks, cloud warehouse) with object storage instead of HDFS.
Move data with DistCp or cloud transfer tools, convert to Parquet/Delta where useful.
Migrate jobs, replace Oozie with Airflow, and test results by comparing row counts and aggregates.
Run in parallel for a period before switching off the old cluster.
What is Master Data Management and what does Informatica MDM do?
Master data is the core business data shared across systems — customers, products, suppliers, employees. MDM creates one trusted "golden record" for each entity.
Informatica MDM loads data from source systems into landing and staging tables, cleanses and standardises it, matches records that refer to the same entity, merges them using trust and survivorship rules, and publishes the best version to other systems. Data stewards handle exceptions in the user interface.
Interview question
What is the difference between match and merge in Informatica MDM?
Match finds records that are likely the same entity using match rules — fuzzy rules (similar names, addresses) or exact rules — and a match score.
Merge combines matched records into one base object record (BVT — best version of truth), using trust settings to choose which source wins for each column.
Matches can be set to auto-merge when the score is high, or queued for manual review by data stewards when uncertain.
Interview question
What are trust and validation rules?
Trust defines how reliable each source system is for each column, and how that trust decays over time — for example CRM is most trusted for email, ERP for billing address.
Validation rules lower the trust of a value that fails a check (for example a phone number in the wrong format), so a better value from another source can win.
Together they decide which value survives in the golden record.
Real-time scenario
After go-live, business users report that obvious duplicate customers are not merging. How do you investigate?
Check the match rules — are the right columns used and is the fuzzy match key populated (tokenisation run)?
Check the data — maybe the names are spelled very differently or key fields are empty after cleansing.
Run match for a sample and look at the match scores; tune match level (conservative/typical/loose) or add a rule.
Check whether the records are marked "consolidated" or excluded from matching.
Test changes on a copy of the data first, because looser rules can cause wrong merges, which are harder to undo.
Real-time scenario
Two different customers were wrongly merged. What do you do?
Use the unmerge function in the data steward UI (or the API) to split the record back to its source cross-references.
Find why it happened — usually a loose fuzzy rule, a shared generic value (like a company phone number), or bad source data.
Tighten the match rule or add an exclusion (for example ignore known generic emails/phones).
Re-run match carefully and check related records that may have been merged the same way.
Real-time scenario
Downstream systems need the golden customer record whenever it changes. How would you deliver it?
Use MDM's publish mechanism — message triggers or the event/publish process — to send changes to a queue (JMS/Kafka) whenever a base object record is updated or merged.
Or expose the data through MDM business entity REST/SOAP services for systems to call on demand.
Include the cross-reference IDs so each system can map the golden record back to its own ID.
Monitor failures and have a replay process for missed messages.
What is the difference between let, const and var?
var — function-scoped, can be redeclared, and is hoisted (initialised as undefined).
let — block-scoped, can be reassigned but not redeclared in the same scope.
const — block-scoped, cannot be reassigned (but objects and arrays declared with const can still be changed inside).
In modern code I use const by default and let only when the value must change; I avoid var.
Interview question
What are props and state in React?
Props are inputs passed from a parent component to a child; the child cannot change them.
State is data owned by the component that can change over time (useState), and changing it re-renders the component.
Example: a ProductList passes each product as props to ProductCard; ProductCard keeps its own "isFavourite" state. When state needs to be shared, I lift it up to a common parent or use context/a state library.
Interview question
What is the difference between REST and GraphQL?
REST — multiple endpoints (/users, /orders/12), each returning a fixed shape; uses HTTP methods and status codes; simple caching.
GraphQL — usually one endpoint; the client asks for exactly the fields it needs in one query, which avoids over-fetching and multiple round trips.
REST is simpler and most common; GraphQL helps when many different clients need different data shapes.
Real-time scenario
A React page becomes slow when showing 10,000 rows. How do you fix it?
Paginate or use infinite scroll from the API instead of loading all rows.
Virtualise the list (react-window/react-virtualized) so only visible rows are rendered.
Avoid unnecessary re-renders — memoise row components (React.memo), stable keys, useMemo/useCallback for heavy calculations and handlers.
Debounce search/filter inputs.
Measure with React DevTools Profiler before and after.
Real-time scenario
How do you implement secure login in a full stack app?
Hash passwords with bcrypt/argon2 — never store plain text.
Use HTTPS everywhere.
After login, issue a session cookie or a short-lived JWT with a refresh token; store tokens in httpOnly, Secure, SameSite cookies, not localStorage.
Validate input on the server, use parameterised queries to prevent SQL injection, and protect against CSRF and XSS.
Add rate limiting and lockout on repeated failed logins, and role checks on every protected API.
Real-time scenario
How would you deploy a React frontend and Node.js/Python backend to production?
Build the frontend (npm run build) and serve the static files from a CDN or object storage (S3 + CloudFront, Azure Static Web Apps).
Containerise the backend with Docker and run it on a managed service (App Service, ECS, Cloud Run) behind a load balancer with HTTPS.
Use a managed database, environment variables/secret manager for config.
Set up CI/CD (GitHub Actions) to test, build and deploy on every merge, with health checks and rollback.
EDI (Electronic Data Interchange) is the standard computer-to-computer exchange of business documents — purchase orders, invoices, shipping notices — between trading partners.
ANSI X12 — the standard mostly used in North America; documents are identified by numbers, for example 850 (Purchase Order), 810 (Invoice), 856 (Advance Ship Notice), 997 (Functional Acknowledgment).
UN/EDIFACT — the international standard used in Europe and other regions; messages like ORDERS, INVOIC, DESADV.
Interview question
Explain the ISA, GS and ST envelopes in X12.
An X12 file has nested envelopes:
ISA/IEA — interchange envelope: sender and receiver IDs, date/time, control number, delimiters.
GS/GE — functional group: groups transactions of the same type (for example all 850s), with its own control number.
ST/SE — transaction set: one business document, like one purchase order, with segment count in SE.
Control numbers must match between the start and end segments, and acknowledgments (997/999) refer to them.
Interview question
What is the difference between a 997 and a 999 acknowledgment?
Both are functional acknowledgments that confirm the receiver got the file and whether it passed syntax checks.
997 — used across most X12 industries; reports accepted/rejected at group and transaction level.
999 — Implementation Acknowledgment, mainly used in healthcare (HIPAA 5010); it also reports implementation guide errors.
They confirm technical receipt only; business acceptance (for example of a PO) is a separate document like the 855.
Real-time scenario
A trading partner says they never received your 810 invoices. How do you troubleshoot?
Check your outbound logs: was the 810 generated and sent, and over which channel (AS2, SFTP, VAN)?
For AS2, check whether you received an MDN (receipt); for VANs, check the VAN tracking.
Check whether you received a 997 from the partner — if not, the file may not have reached their translator.
Verify partner IDs (ISA06/ISA08), qualifiers and certificates/endpoints have not changed.
Resend after fixing, avoiding duplicate control numbers, and confirm with the partner.
Real-time scenario
How do you onboard a new trading partner for 850 purchase orders?
Get their implementation guide (which segments and elements they send) and connection details (AS2/SFTP/VAN, IDs, certificates).
Set up the trading partner profile in the EDI tool (IDs, qualifiers, envelopes).
Build or adjust the map from their 850 to your ERP format (for example SAP IDoc ORDERS05).
Test with sample files: connectivity, 997 exchange, mapping results in the ERP.
Go live with monitoring and an error alert process.
Real-time scenario
Partner sends a new optional segment in the 850 and your map fails. How do you handle it?
Check the translator error to see which segment failed validation.
Compare with the partner's implementation guide version — they may have moved to a new version.
Update the map/standard definition to allow the segment (or ignore it if not needed), test with their sample file.
Reprocess the failed files after the fix.
Ask partners to notify changes in advance and keep versioned maps so a rollback is possible.
What are the main organisational units in SAP FI and CO?
Company — the legal group entity for consolidation.
Company Code — the smallest unit with its own balance sheet and P&L (FI).
Chart of Accounts — list of GL accounts used by company codes.
Business Area / Segment — for internal reporting.
Controlling Area — the CO unit; one or more company codes are assigned to it.
Cost Centers, Profit Centers and Internal Orders sit under the controlling area for cost tracking.
Interview question
What is the difference between a reconciliation account and a normal GL account?
A reconciliation account is a GL account linked to a subledger — customers (AR), vendors (AP) or assets. You cannot post to it directly; postings come automatically when you post to a customer, vendor or asset, so the GL always matches the subledger.
A normal GL account (for example bank charges, rent) is posted directly. The reconciliation account is assigned in the customer/vendor master or asset class.
Interview question
What is automatic account determination in MM-FI integration (OBYC)?
When a goods movement or invoice is posted in MM, SAP automatically decides which GL accounts to post to, using transaction OBYC. It uses the transaction key (for example BSX for inventory, WRX for GR/IR clearing, GBB for offsetting entries), valuation grouping code, valuation class (from the material master) and account modifier.
Example: goods receipt for a PO posts Debit Inventory (BSX) and Credit GR/IR (WRX).
Real-time scenario
During month-end closing, the GR/IR account shows old open items. What do you do?
Analyse open items with MB5S (GR/IR balances) to see if goods were received without invoice or invoiced without receipt.
Coordinate with purchasing and AP to post missing invoices or goods receipts, or clear wrong ones.
Run automatic clearing (F.13) for items that match.
For the balance sheet, run GR/IR regrouping (F.19 / FAGLF101 in S/4HANA) to show the correct liability or asset.
Document any long-pending items for audit.
Real-time scenario
The automatic payment run (F110) did not pick up an invoice that is due. How do you troubleshoot?
Check the vendor master: payment methods, bank details, payment block.
Check the invoice: payment block, due date (baseline date + terms), payment method.
Check the F110 parameters: company code, payment methods, next run date, vendor range.
Read the F110 exception list/log — it shows why each item was not paid.
Check the payment program configuration (FBZP) for that payment method and house bank.
Real-time scenario
Management wants to see marketing costs by campaign. How would you set it up in CO?
Create an Internal Order for each campaign (order type for marketing) with a responsible cost center.
Post marketing expenses to the internal order instead of only the cost center.
Use budgets with availability control if spending limits are needed.
At period end, settle the orders to the marketing cost center (or to profitability analysis) using settlement rules.
Report with order reports (for example S_ALR_87012993) or Fiori analytics.
In S/4HANA Finance, the table ACDOCA stores all accounting line items — GL, CO, asset accounting and material ledger — in one place.
Benefits: one source of truth (no reconciliation between FI and CO), real-time reporting, more dimensions on each line, and faster closing. Old tables like BSEG, COEP and GLT0 still exist or are available as compatibility views for older programs.
Interview question
Why is Business Partner mandatory in S/4HANA?
In S/4HANA, customers and vendors are maintained through the Business Partner (transaction BP), with customer and vendor roles. Customer/Vendor Integration (CVI) keeps the classic customer and vendor tables in sync.
Benefits: one central master for a party that can be both customer and vendor, shared addresses and relationships, and a single maintenance transaction. During conversion, CVI must be completed before moving to S/4HANA.
Interview question
What changed in Asset Accounting in S/4HANA?
S/4HANA uses New Asset Accounting:
Asset postings go to the Universal Journal in real time for all depreciation areas.
Parallel valuation (for example local GAAP and IFRS) is handled with parallel ledgers or accounts.
Periodic posting of APC values (old ASKB) is not needed.
Asset reconciliation with GL is automatic.
Configuration of depreciation areas must be consistent with the ledger setup.
Real-time scenario
Your company is moving from ECC to S/4HANA Finance. What finance-specific checks do you do?
Complete Customer/Vendor Integration to Business Partner.
Run the finance readiness checks and data consistency checks (FIN_CORR reconciliations) before conversion.
Migrate to New Asset Accounting and check the ledger and currency setup.
Plan the migration of balances and open items, and reconcile after conversion (GL, AR, AP, assets).
Retest custom reports that used old tables like BSEG/GLT0, and train users on Fiori apps.
Real-time scenario
Managers want a real-time P&L by profit center and segment. How do you provide it in S/4HANA?
Make sure profit center and segment are derived on every posting (document splitting, substitution or derivation from cost center/material).
Use Fiori apps on the Universal Journal like "Display Financial Statement" or "Profit Center P&L"/Trial Balance with the needed dimensions.
For custom layouts, use embedded analytics: CDS views and Analysis for Office or SAP Analytics Cloud.
Because everything is in ACDOCA, the data is real time with no separate reconciliation.
Real-time scenario
A balance sheet by profit center does not balance. What could be the reason?
Document splitting is not active, or the splitting characteristics (profit center, segment) are not set as zero-balance.
Some postings came without a profit center — check the default/dummy profit center and derivation rules.
Business transaction variants or item categories are wrongly assigned for some document types.
Run the splitting check for affected documents and correct the configuration; for history, use the provided correction tools after testing.
Central Finance is an S/4HANA system that receives accounting documents in real time from many source systems (SAP ECC, older SAP, or non-SAP) and stores them in one Universal Journal.
It gives one central view of finance across the group without changing the source systems first — often used as a first step towards S/4HANA, for shared services, central reporting and central payments.
Interview question
How are documents replicated to Central Finance?
SAP Landscape Transformation (SLT) captures changes in the source systems' tables and sends them to the Central Finance system. There, the accounting interface reposts them as Central Finance documents.
Mapping (MDG mapping or key mapping) converts source values like company codes, GL accounts and cost centers to the central values. Errors are handled in the Application Interface Framework (AIF).
Interview question
Why is master data mapping important in Central Finance?
Source systems often use different codes — different GL account numbers, cost centers or customer IDs. Central Finance needs one harmonised set.
Mapping tables (value mapping and key mapping, often managed with SAP MDG) translate each source value to the central value. Without correct mapping, documents fail in replication or post to wrong accounts, so mapping is a big part of a Central Finance project.
Real-time scenario
Many documents are failing during replication to Central Finance. How do you troubleshoot?
Check the errors in AIF (Application Interface Framework) — they show the reason per document.
Common causes: missing mapping for an account or cost center, master data not created in the central system, closed posting period, or configuration differences.
Fix the root cause (add mapping, create master data, open period) and reprocess from AIF.
Monitor SLT for replication delays.
Build a daily error report so the team clears errors before month-end.
Real-time scenario
How do you plan the initial load of historical data to Central Finance?
Decide what to load: balances, open items, and how many years of line items.
Prepare mapping and master data first, and lock the configuration.
Run the initial load in a test system, compare totals with the source (by company code, account, period) and fix errors.
Plan the production load with a cut-off date, then switch on real-time replication for new postings.
Reconcile again after go-live.
Real-time scenario
The CFO wants one group-wide P&L from five different ERP systems. How does Central Finance help?
All five systems replicate their accounting documents to Central Finance, mapped to one chart of accounts and one set of profit centers.
Reports run on the central Universal Journal, so the CFO gets one real-time view across all systems.
Drill-down to the original source document is possible.
This can be extended with central payments, central controlling and group reporting.
Source determination — info records, source list, RFQ/quotations.
Purchase Order (ME21N) — sent to the vendor.
Goods Receipt (MIGO, movement type 101) — updates stock and posts inventory/GR-IR.
Invoice Verification (MIRO) — three-way match of PO, GR and invoice.
Payment — done in FI (F110 or manual).
Interview question
What are common movement types in SAP MM?
101 — goods receipt for PO.
102 — reversal of 101.
122 — return delivery to vendor.
201 — goods issue to a cost center.
261 — goods issue to a production order.
301/311 — transfer posting between plants / storage locations.
551 — scrapping.
561 — initial stock entry.
Each movement type controls the stock and accounting updates.
Interview question
How does pricing work in a purchase order?
PO pricing uses the condition technique: a calculation schema (pricing procedure) is determined from the purchasing organisation schema group and vendor schema group. The schema lists condition types like gross price (PB00), discounts and freight. Values come from condition records — usually from the purchasing info record. The final net price is calculated step by step in the procedure.
Real-time scenario
MIRO shows a price variance and the invoice is blocked for payment. What do you do?
Check the variance: PO price vs invoice price, and quantity received vs invoiced.
Confirm with the buyer whether the vendor's price is correct (maybe a price change not updated in the PO).
If correct, release the invoice with MRBR (or adjust the PO and re-check); if wrong, ask the vendor for a credit memo.
Check the tolerance keys (OMR6) if many small variances block invoices unnecessarily.
Real-time scenario
Stock in SAP does not match the physical warehouse count. How do you correct it?
Run a physical inventory: create the document (MI01), enter counts (MI04), check differences (MI20), post differences (MI07).
Investigate big differences before posting — missing goods receipts, issues not posted, wrong storage location transfers.
Use cycle counting for high-value items to avoid big year-end differences.
Posting differences creates accounting entries to the inventory difference account.
Real-time scenario
Purchase orders above 1 lakh must be approved by the purchase manager. How do you configure it?
Set up a PO release strategy with classification:
Create characteristics (for example net value CEKKO-GNETW, purchasing group, document type) and a class.
Define release groups, release codes (purchase manager), release indicators and the strategy.
Maintain the classification values: net value > 100,000 INR.
The approver releases with ME29N or the Fiori app; until then the PO cannot be output to the vendor.
In S/4HANA, flexible workflow for purchase orders is the newer alternative.
Sales Order — VA01, with availability check, pricing and credit check.
Delivery — VL01N, picking, packing.
Post Goods Issue — reduces stock and posts cost of goods sold.
Billing — VF01, creates the invoice and the accounting document in FI.
Payment from the customer — posted in FI (F-28 or incoming payments).
Interview question
How does the condition technique work in SD pricing?
Pricing procedure determination uses sales area + customer pricing procedure + document pricing procedure. The pricing procedure lists condition types (PR00 price, K007 discount, MWST/tax, freight). Each condition type has an access sequence that searches condition tables in order (for example customer-material, then material only) to find condition records. The first valid record found is used.
Interview question
What is copy control in SD?
Copy control defines how data is copied from one document to the next — quotation to order (VTAA), order to delivery (VTLA), delivery or order to billing (VTFL/VTFA). It controls item categories, copying requirements (routines that must be met), data transfer routines and whether pricing is copied or redetermined.
Real-time scenario
A sales order shows "Pricing error: mandatory condition PR00 is missing". What do you check?
Check whether a condition record exists for PR00 for that customer/material/sales area and validity date (VK13).
Check the access sequence of PR00 — maybe the record is maintained with a key combination not in the access sequence.
Check the sales area and pricing procedure determined in the order (Analysis in the condition tab shows why each access failed).
Create or correct the record and use "Update pricing" in the order.
Real-time scenario
A customer's order is blocked for credit. How does it work and how do you release it?
Credit management checks the order value plus open orders, deliveries and receivables against the customer's credit limit (in S/4HANA via FSCM Credit Management with the business partner credit profile).
If the limit is exceeded, the order is blocked for delivery.
An authorised credit controller reviews and releases it with VKM1/VKM3 (or the Fiori app "Manage Documented Credit Decisions").
If needed, the credit limit is increased in the BP credit segment.
Real-time scenario
How do you process a customer return in SAP SD?
Create a return order (order type RE) with reference to the billing document.
Create a return delivery and post goods receipt (movement type 651 to returns/blocked stock).
Inspect and move the stock to unrestricted or scrap as needed.
Create a credit memo (billing type RE) to refund the customer.
Optionally use Advanced Returns Management for more complex return processes.
What is SAP EWM and how is it different from SAP WM?
SAP EWM (Extended Warehouse Management) is SAP's warehouse solution for complex, high-volume warehouses. Classic WM (Warehouse Management) is no longer the strategic option in S/4HANA, and customers are expected to move to EWM.
EWM supports waves, labour management, slotting, cross-docking, yard management, value-added services and material flow systems (automation).
EWM can run embedded in S/4HANA or as decentralised EWM on a separate system.
Stock in EWM is managed in storage bins and handling units, with warehouse tasks and warehouse orders driving the work.
Interview question
Explain the EWM organisational structure.
Warehouse number — the whole physical warehouse, linked to ERP plant and storage location.
Storage type — an area with the same putaway/picking method (high rack, bulk, fixed bin, GR zone).
Storage section — a group of bins in a storage type (fast movers, heavy goods).
Storage bin — the smallest address where stock is kept.
Activity area — a logical grouping of bins used for picking and physical inventory.
Doors and staging areas — where goods are loaded and unloaded.
Interview question
What is the difference between a warehouse task and a warehouse order?
A warehouse task (WT) is the instruction to move a quantity of a product from a source bin to a destination bin — putaway, picking or internal movement.
A warehouse order (WO) groups warehouse tasks into one work package for one worker, built using warehouse order creation rules (by activity area, weight, number of items).
Workers usually execute warehouse orders on RF devices; confirming the WO confirms its tasks and updates stock.
Real-time scenario
Goods are received but the putaway warehouse task is not created. How do you troubleshoot?
Check the inbound delivery in EWM — status, whether the goods receipt was posted, and the process type.
Check the putaway strategy: storage type search sequence, and whether the product master has the correct putaway control indicator.
Check whether suitable empty bins exist with the right bin type and capacity.
Look at the application log and the warehouse monitor for errors.
Check the warehouse process type and whether automatic WT creation is switched on.
Fix the master data or configuration, then create the task again from the delivery.
Real-time scenario
During picking the RF user finds the bin empty although the system shows stock. What do you do?
The user reports a bin denial / exception code on the RF device instead of confirming the task.
The exception code triggers a follow-up: the system searches another bin for the remaining quantity and creates a new task.
The original bin is blocked or flagged for a physical inventory count.
After the count, the difference is posted through the difference analyser so EWM and S/4HANA stock match.
Later, check why it happened — missed confirmation, wrong putaway or theft — and fix the process.
Real-time scenario
Outbound orders must be shipped by truck at 4 PM. How do you plan picking in EWM?
Use wave management: a wave template with release times so that all deliveries for the 4 PM truck are grouped in one wave.
Release the wave early enough to allow picking, packing and staging before loading.
Use warehouse order creation rules to split picking work efficiently by activity area.
Stage goods in the staging area for the right door, then load and post goods issue.
Monitor progress in the warehouse monitor and handle delays by re-prioritising.
What are the key changes in logistics when moving from ECC to S/4HANA?
Business Partner replaces separate customer and vendor masters.
MATDOC is the single table for material documents; aggregate tables are calculated on the fly.
Material number length can be extended to 40 characters.
Credit management moves to SAP Credit Management (FIN-FSCM).
MRP Live runs MRP in HANA; output management can use the new BRF+ based framework.
Fiori apps replace many classic transactions for daily work.
Interview question
Why is Business Partner mandatory in S/4HANA and how are customers and vendors linked to it?
In S/4HANA the Business Partner (transaction BP) is the single entry point for customers and suppliers. It holds general data once (name, address, bank) and uses roles for customer (FI and SD) and supplier (FI and MM) data.
Customer/Vendor Integration (CVI) keeps the classic customer and vendor tables in sync with the BP. During conversion, CVI must be set up and all customers and vendors converted to Business Partners before the system move.
Interview question
What is MRP Live?
MRP Live (transaction MD01N) is the S/4HANA MRP run optimised for HANA. Most of the planning logic runs in the database, so it is much faster than classic MRP.
It plans across plants and handles stock transfers.
It can be scheduled as a background job or used through Fiori apps like Monitor Material Coverage.
Classic MD01 still exists but MRP Live is the recommended way.
Real-time scenario
During an ECC to S/4HANA conversion, simplification checks report open items in logistics. How do you handle them?
Run the Simplification Item Check and review each item with the business and functional teams.
Explain the SAP NetWeaver ABAP system architecture.
Presentation layer — SAP GUI, Fiori launchpad or browser.
Application layer — application servers with the dispatcher and work processes; the ABAP Central Services (ASCS) instance holds the message server and enqueue server.
Database layer — one database (HANA for S/4HANA) shared by all application servers.
Users log on via the message server, which balances load across application servers using logon groups.
Interview question
What are the different types of work processes?
Dialog (DIA) — interactive user requests.
Background (BTC) — scheduled jobs.
Update (UPD/UP2) — database updates in V1 and V2 tasks.
Enqueue (ENQ) — lock management (runs in ASCS in modern systems).
Spool (SPO) — print requests.
Monitor them in SM50 (local) and SM66 (all servers); parameters like rdisp/wp_no_dia control the count.
Interview question
How does the transport system work?
The Transport Management System (STMS) moves changes through the landscape — usually DEV → QAS → PRD.
Changes are recorded in transport requests (SE09/SE10) in DEV.
Releasing a request exports the data and cofiles to the transport directory (/usr/sap/trans).
The request appears in the import queue of QAS, is imported and tested, then imported into PRD.
Transport routes and layers are defined in STMS on the domain controller.
Real-time scenario
Users say the system is very slow since morning. What do you check?
SM50/SM66 — are all dialog work processes busy? Which programs are running long?
ST22 — any dumps in large numbers; SM21 — system log errors.
ST06/OS monitor — CPU, memory and paging on servers.
ST04 or HANA Cockpit/DBACOCKPIT — expensive SQL statements, locks, memory.
SM12 — old lock entries; SM13 — update errors.
SM37 — heavy background jobs running in business hours.
Fix the immediate cause (stop a runaway job, add resources), then do root cause analysis.
Real-time scenario
A transport fails in QAS with return code 8. What do you do?
Open the transport log in STMS and find the exact error — usually a syntax error, missing object or activation error.
Check whether a dependent transport was not imported first (wrong sequence).
Inform the developer with the log details; they fix the object in DEV and release a new transport.
Import the dependent request first, then re-import.
Return code 4 is a warning; 8 means errors that must be fixed; 12 or higher is a serious import problem.
Real-time scenario
The business asks you to refresh QAS with production data. What are the main steps?
Plan downtime and inform users; export QAS-specific settings — RFC destinations, STMS config, users, printers, licences, logical systems.
Take a backup of PRD (or use a recent backup) and restore it on QAS.
Run post-copy steps: rename the system, change logical system names (BDLS), adjust RFCs, reset STMS, suspend jobs, apply QAS licence.
Scramble or mask sensitive data if needed.
Validate with the functional teams, then release the system.
What is the difference between single, composite and derived roles?
Single role — has transactions/apps and authorisation objects; generates a profile.
Composite role — a container of single roles assigned together; no authorisations of its own.
Derived role — inherits menu and authorisations from a parent (master) role but has different organisational values, such as company code or plant.
Derived roles keep design consistent across company codes and reduce maintenance.
Interview question
A user gets "no authorisation". How do you analyse it?
SU53 — shows the last failed authorisation check for that user: object, field and values.
ST01 or STAUTHTRACE — trace all authorisation checks during the user's activity, useful when SU53 is not enough.
SUIM — find which roles contain the needed object values.
Then add the right value to the correct role (following the role design and approval process), never by giving broad access like SAP_ALL.
Interview question
What are the main components of SAP GRC Access Control?
Access Risk Analysis (ARA) — checks users and roles for segregation of duties (SoD) risks against a rule set.
Access Request Management (ARM) — request and approve access through workflow.
Emergency Access Management (EAM) — firefighter IDs for urgent tasks, with full logging and review.
Business Role Management (BRM) — role design and approval lifecycle.
User Access Review (UAR) — periodic certification of user access by managers.
Real-time scenario
Audit finds a user who can create a vendor and also pay that vendor. What do you do?
This is a segregation of duties (SoD) risk. Run Access Risk Analysis to confirm the conflicting functions and roles.
Talk to the business owner: can the access be split between two people?
If yes, remove the conflicting role and re-assign properly.
If not (small team), apply a mitigating control — for example a monthly review of vendor master changes and payments — and assign it in GRC with an owner.
Document the decision for the auditors.
Real-time scenario
Production support urgently needs change access in PRD at night. How is it handled?
Use Emergency Access Management: the support person logs in with an assigned firefighter ID through GRC.
They enter the reason and ticket number; all actions are logged.
After the session, the firefighter controller gets a log and reviews the activity.
Firefighter IDs are given only to approved people and are reviewed regularly.
This gives fast access without permanently giving powerful roles.
Real-time scenario
A user can see a Fiori tile but the app shows an error when opened. What do you check?
Front-end: the role has the catalog and group/space with the tile, and target mapping is correct.
OData service: the service is activated in /IWFND/MAINT_SERVICE and the user has S_SERVICE authorisation for it in the front-end role.
Back-end: the user has the business authorisations (for example company code values) in the back-end role.
Check /IWFND/ERROR_LOG and SU53 on both front-end and back-end.
Fix the missing authorisation in the correct role.
What happened to SAP Leonardo and where are its capabilities now?
SAP Leonardo was SAP's brand (around 2017) for intelligent technologies — machine learning, IoT, blockchain, analytics and design thinking. SAP retired the Leonardo brand around 2019.
Its capabilities moved into SAP Business Technology Platform (SAP BTP): AI services, SAP AI Core and AI Launchpad, IoT services, SAP Analytics Cloud and integration services. Today, in interviews, explain your skills in terms of SAP BTP and SAP Business AI.
Interview question
What is SAP BTP and what are its main areas?
SAP Business Technology Platform brings SAP's cloud technology together:
Application development — SAP Build, Cloud Foundry and Kyma runtimes, ABAP environment, CAP.
Integration — SAP Integration Suite (CPI, API Management, Event Mesh).
Data and analytics — SAP Datasphere, SAP HANA Cloud, SAP Analytics Cloud.
AI — SAP AI Core, AI Launchpad, Generative AI Hub, Joule.
Accounts are organised as global account → subaccounts → spaces/namespaces.
Interview question
What is SAP AI Core and AI Launchpad?
SAP AI Core runs AI workloads on BTP — training and serving models in containers (based on Kubernetes), and gives access to large language models through the Generative AI Hub.
SAP AI Launchpad is the user interface to manage AI scenarios, deployments, connections and to test prompts.
Together they let you add AI to SAP processes with governance and connection to business data.
Real-time scenario
A manufacturing client wants to predict machine failures using sensor data with SAP. How would you design it?
Collect sensor data via IoT gateway/edge into the cloud (SAP or hyperscaler IoT service).
Store and model data in SAP HANA Cloud or Datasphere, combined with maintenance history from S/4HANA.
Train a predictive model (for example failure classification) in SAP AI Core.
Deploy the model and call it from a BTP app or workflow.
When failure risk is high, create a maintenance notification in S/4HANA automatically.
Show results in SAP Analytics Cloud dashboards.
Real-time scenario
The business wants an AI assistant to answer questions from internal SAP documents. How do you approach it on BTP?
Use the Generative AI Hub in SAP AI Core to access an LLM with enterprise controls.
Build retrieval-augmented generation (RAG): split documents into chunks, create embeddings and store them in the SAP HANA Cloud vector engine.
On each question, search similar chunks and send them with the question to the LLM.
Build the app with CAP or SAP Build and secure it with BTP roles.
Add guardrails: data masking, logging, and testing answers for accuracy.
Real-time scenario
A customer still has an old SAP Leonardo ML service. What do you recommend?
Find out which services they use — older Leonardo ML foundation services have been retired.
Map each use case to the current option: SAP AI Core models, SAP AI Business Services such as Document Information Extraction, or Generative AI Hub.
Re-build and test in a BTP subaccount, compare results with the old service.
Plan cutover and update integrations.
This avoids running unsupported services and gives access to newer AI features.
What is SAP CPI and how is it related to SAP Integration Suite?
SAP Cloud Platform Integration (CPI), now called Cloud Integration, is the iPaaS service for building integration flows (iFlows) between SAP and non-SAP systems in the cloud.
It is part of SAP Integration Suite on BTP, which also includes API Management, Open Connectors, Integration Advisor, Event Mesh and prepackaged integration content.
Interview question
Name common adapters used in CPI.
SOAP, HTTP/HTTPS and REST (via HTTP) for web services.
IDoc and RFC for SAP ECC/S/4HANA.
OData for SAP S/4HANA and SuccessFactors APIs.
SFTP and Mail for files and emails.
SuccessFactors, Ariba and JMS adapters.
Process Direct to call one iFlow from another.
On-premise systems are reached through the SAP Cloud Connector.
Interview question
When do you use a Groovy script in CPI?
When standard steps (content modifier, message mapping, converters) are not enough — for complex logic, dynamic headers or properties, custom logging or payload manipulation.
What is code-to-data (code pushdown) in ABAP on HANA?
Instead of reading large data into the application server and processing it in ABAP loops, the calculation is moved to the HANA database.
Tools for this:
Open SQL / ABAP SQL with joins, aggregates, CASE and expressions.
CDS views for reusable data models.
AMDP (ABAP Managed Database Procedures) for complex SQLScript logic.
This reduces data transfer and makes reports much faster.
Interview question
What are CDS views and what are annotations used for?
Core Data Services (CDS) views define data models in the database using SQL-like syntax in ADT (Eclipse).
They support joins, associations, calculations, parameters and access control (DCL).
Annotations add meaning: @OData.publish or service definitions, @UI annotations for Fiori elements, @Analytics for reporting, @AccessControl for authorisation checks.
In S/4HANA, the newer CDS view entities (define view entity) are recommended.
Interview question
What is the ABAP RESTful Application Programming Model (RAP)?
RAP is the modern way to build Fiori apps and APIs in S/4HANA and the ABAP Cloud environment.
Data model — CDS view entities.
Behaviour definition — create, update, delete, actions, validations and determinations.
Behaviour implementation — ABAP class with the logic.
Service definition and service binding — expose as OData V2/V4 UI or API.
It supports managed and unmanaged scenarios and draft handling.
Real-time scenario
An old custom report takes 30 minutes after moving to S/4HANA. How do you fix it?
Run SAT (runtime analysis) and ST05 (SQL trace) to find the slow parts.
Look for SELECT inside LOOP, SELECT * and FOR ALL ENTRIES on huge tables.
Replace with joins or a CDS view so HANA does the aggregation.
Use the SQL Monitor (SQLM) and the ABAP Test Cockpit with HANA checks to find more issues.
Read only needed columns, and use proper WHERE conditions.
Re-test with production-like data and compare runtime.
Real-time scenario
Before an S/4HANA conversion, how do you check custom code?
Use the Custom Code Migration app or ATC with the S/4HANA readiness check variant.
It finds code using removed tables or changed data models (like MATDOC, BP, field length extensions).
Remove unused code first — check usage with SCMON/UPL data.
Fix findings: replace old tables with new ones or released APIs, adjust field lengths, fix ORDER BY assumptions.
Re-run ATC until it is clean.
Real-time scenario
The customer wants a "clean core" and asks you to add a custom field and logic to sales orders. How do you do it?
Prefer key-user extensibility: add the custom field with the Custom Fields app, and logic with Custom Logic (BAdI) using released APIs.
For bigger logic, use developer extensibility with ABAP Cloud — only released objects, in a separate software component.
For side-by-side needs, build on SAP BTP and call S/4HANA APIs.
Avoid modifications to SAP standard code.
This keeps upgrades simple and follows SAP's clean core guidance.