A classmate holds up a photo of the access policy. The heading is neat. The date is this year. She says leaver access is fine. You ask a smaller question. Which person who left in March did you test? She has not tested one. The policy is a rule. It is not evidence that the rule ran.
That gap is fieldwork. Fieldwork is detective work. You already know the question, because you wrote it in the plan. Now you collect proof. Then you write the proof down so another auditor can rerun it without calling you.
In Part 1, What Is an IT and IS Audit?, you met the four phases, and you saw a leaver test recorded as working paper A-12. In Part 2, Phase 1 of an IS Audit, you learned how to get authority, understand BBC Bank, score risk, and write the plan before anyone opens a sample. This is part 3 of 5 in Understanding IT and IS Audit. The plan said what and why. The audit program, which you start now, says how. Fieldwork is you doing that how, and keeping the record.
You are close to the BBC Bank capstone. The pull is to click until something looks wrong. A click with no objective is a tour. An audit is a test you can explain on one page.
Every banking picture in this article is an illustration for class. It is not a template to copy onto a live bank, and it is not a report on a named bank in town. On the capstone, use the system names, the dates, and the policy clauses in your case pack. If the pack is quiet, label the choice as a teaching choice, and say that you did so.
What Fieldwork Is
Fieldwork is Phase 2. It is the part of the audit where you carry out the tests and gather the evidence. Planning asked, "What must we be able to conclude?" Fieldwork asks, "What did we actually find, and where is the proof?"
Part 2 kept two documents apart. The audit plan says what you will cover, and why. The audit program says how you will test. You write the program at the start of fieldwork, from the plan, not from whatever file happens to open. A clever test that does not serve the objective does not belong in the program. A risk you scored at 20, with no test that touches it, means the program dropped a promise the plan made.
A program line is short. It names the objective, the control, the test, the population, and the working paper reference. The detail of what you did lives on the working paper, not in a chat. Part 1 already gave you the rule. If it is not in the papers, it did not happen.
Fieldwork is not the week you fix the bank. You assess the control. You do not redesign it. If you disable the old accounts yourself so the list looks clean, you have left the audit and joined operations. Independence, from Part 1, is still in force. You may tell management what you found. You may not quietly repair it and then give assurance on your own repair.
On a normal day the work has a rhythm.
Read the objective and the criteria again. Say them out loud before you open a file.
Name the population. Write the source, the period, and the count.
Choose the sample, and write why that choice fits the risk.
Do the test. Prefer a check you can rerun over a story someone tells you.
Keep the evidence with a file name, a date, and the rows that matter.
Write the conclusion on the working paper the same day. Do not trust tomorrow's memory.
You will talk to people. Talking is how you learn where the button is. Talking is not how you finish. A sentence that begins "the officer said" needs a second sentence that begins "the log shows".
Bring the Part 2 numbers with you, but do not redo the whole plan. Agency reversals scored 20 out of 25. Materiality for a posting error was UGX 1,000,000. Performance materiality was UGX 600,000. The period in the worked example was 1 April 2026 to 30 September 2026. Those are teaching choices from the plan. Fieldwork uses them. It does not invent a new ruler after the sample looks awkward.
Two Families of Controls
Before you test, sort the control. Auditors talk about two families. If you mix them up, you will test a fee formula and think you have tested the whole bank.
IT general controls
IT general controls, often shortened to ITGCs, are the controls around the environment that many systems share. They are not about one payment. They are about whether the systems that process payments can be trusted at all. In this course, four ITGCs come up again and again.
Access. The right people can get in. Leavers cannot. Powerful accounts are few, named, and reviewed.
Change. Nobody alters a live system without a request, a test, an approval, and a record.
IT operations. Jobs run, failures get attention, and the bank can recover. Backups and a tested restore sit here.
The split of powerful duties. The person who builds a change is not the only person who can move it live.
Part 1 already showed these in the night window. A leaver, a day-shift teller, or a shared ID such as NIGHT01 on the reversal screen is an access failure. A reversal limit changed on a Thursday, called temporary, and never put back, is a change failure. A disaster recovery document that has never been restored onto a clean machine is an operations failure. The policy in the drawer is not the control. The tested restore is.
Application controls
Application controls sit inside one business process. They try to make that process complete, accurate, and authorised. Hall and the CISA material group them by where they act. Use that grouping. It stops you from calling every check "a control" and then testing none of them.
Input controls stop bad data getting in. A required field, a valid account number, a date that cannot be in the future, and a warning when the amount is blank are input controls. So is a limit that refuses a teller payment above a ceiling, or routes it to a second person.
Processing controls check the work while the system calculates. A fee taken from the tariff, a tax worked from that fee, and a debit that must have a matching credit are processing controls. When you recompute the fee yourself, you are testing this family.
Output controls check the result after processing. A morning reconciliation of agent float, a review of a suspense report, and a check that yesterday's file reached finance are output controls. They do not stop the bad item at the door. They should catch it before the day is treated as closed.
Two application controls deserve their own names, because you will meet them on the capstone in the first week.
Maker-checker means the person who starts a transaction is not the person who approves it. The maker creates. The checker approves. If the system lets one user do both, the control is not there, even when the officer is honest. Part 2 used this on agency reversals. The objective was to determine whether reversals were approved by someone other than the initiator, and logged to a named user.
A limit is a ceiling the system knows. Below the line, one officer may proceed. Above the line, the system blocks the item or demands a stronger approver. A limit that lives only in a speech at induction is not a control. A limit in a parameter, enforced by the system, is a control. You test the parameter and the items that crossed it.
Why you test general controls first
Test ITGCs before you lean on application controls. The reason is practical. Application controls live in software and in settings. If access is loose, a leaver can still approve payments. If change control is loose, someone can edit the fee table tonight and leave no ticket. Your recomputation of last month's fees can be perfect, and still say nothing about tomorrow, because the program that charges the fee is not protected.
There is a second reason, and it joins Part 2. Control risk is the risk that the controls fail or do not exist. If the general controls are weak, control risk on the reports themselves goes up. A clean application test, run on a report that anyone could have edited, is weaker evidence than it looks. You then need more substantive testing, so detection risk stays low enough. High and high still means test harder. Fieldwork is where that sentence becomes hours and sample size.
You do not need a perfect ITGC opinion on the whole bank before you look at one payment. You do need a view on the general controls that protect this process. For agency reversals, that means access to the reversal screen, change control on the reversal limit, and the backup that would let the bank recover the log. If those are soft, say so on a working paper, and do not claim the maker-checker passed as if the ground under it were solid.
How to Test the Controls You Will Actually Meet
This is the section to use on a lab afternoon. Each control below has the same shape, so you can copy the shape and not the story. The shape is the method. The bank details are illustrations.
For each one you will see what the control is, the objective in one sentence, whether the test is compliance or substantive, the population, a sample idea, the steps, the evidence you keep, and what an exception looks like. Do them in this order on the day: write the objective, get the two files you need, test, then write. If a file is refused, the refusal goes on the paper. You do not fill the gap with a confident sentence.
Leaver access removal
What the control is. This is an ITGC. When someone leaves BBC Bank, their access to the core banking system is removed within one working day of the last working day. Friday leavers are off by Monday. A shared mailbox is not a substitute for disabling the user. Part 1 used this picture because a leaver who can still sign in is a stranger with a key.
Objective. To determine whether access for staff who left in March was removed within one working day.
Kind of test. Compliance. You are asking whether the control operated as the policy requires. You are not yet asking whether those people moved money after they left. If you find open accounts, a later substantive look at their logins is wise. The first test is still the compliance test.
Population. Every staff member whose last working day fell in March. In the teaching file that list has 25 people. HR owns the list. IT does not get to shrink it by saying some leavers were only temps.
Sample. All 25. The population is small, and a missed leaver can see customer data. When the list is this short, sampling is a way to hide. Test the lot. If your case pack has 400 leavers in a year, test a sample, and also test every leaver who had a privileged role.
Steps.
Ask HR for the March leaver list: staff id, name, department, last working day. Save it as hr_leavers_march.csv. Do not retype it.
Ask for a core user extract dated at the end of March, or later: staff id, user id, status, disabled date. Save it as core_users_extract.csv. Watch the extract being run if you can. A file that arrives with no date in the name is harder to trust.
Match on staff id. Use XLOOKUP in Excel, or the small pandas merge later in this article. An inner match that is still Active is the problem set.
For each disabled account, count working days between the last working day and the disabled date. One working day is the rule. A Friday exit may show a Monday disable. That can pass. A Friday exit still open the next Friday does not.
Look for the leaver who has no row on the user extract. That may mean they never had access, or it may mean the id does not match. Do not call it a pass until HR or IT shows you why the ids differ.
Evidence you keep. Both exports, the match file, the policy clause that says one working day, and a note of who ran the extract and on what date. In the worked paper later, the exception rows are 14 and 22.
What an exception looks like. Staff 1044 left on 12 March. On the 31 March extract the core account is still Active. Staff 1182 left on 6 March and was disabled on 20 March. That is more than one working day. Both are exceptions. A manager's note that says "we know them, it is fine" is not a pass. It is a comment you attach to the exception.
Same day. Send the two requests this morning. Match them this afternoon. Write WP-A12 before you go home, even if the conclusion is "extract not yet received".
Privileged and admin access review
What the control is. This is an ITGC. Privileged access means a user can do more than a normal officer: create users, change parameters, read any account, or approve above every limit. The control is a periodic review. A manager who does not hold that admin role looks at the list, signs it, and removes anyone who should not be there. A review that the admins sign for themselves is a weaker review.
Objective. To determine whether privileged access on the core banking system was reviewed in the quarter, and whether access that was no longer needed was removed.
Kind of test. Compliance. You are testing whether the review happened and whether it had an effect. You are not, in this test, tracing a stolen payment.
Population. Every account with an admin, super-user, or parameter role on the core at quarter end, plus the review record for that quarter. Include the switch if reversals can be approved there. Include shared ids. NIGHT01 is in this population even if nobody calls it a person.
Sample. If the privileged list has fewer than 40 accounts, test all of them. If it is longer, test every shared id and every domain-level admin, then sample the rest. Do not sample only the accounts with tidy names.
Steps.
Export the role list: user id, name, role, last login, status. Save it with the system name and the date.
Get the signed review for the quarter. You want the date, the reviewer's name, and the pages that list the users. A calendar invite called "access review" is not the review.
Compare the export to job titles. A teller with a parameter role is a question. A developer with day-to-day admin on the live core is a question.
For each user the review said to remove, check the current extract. Removed on paper and still active in the system is an exception.
Search the list for generic names: ADMIN, NIGHT01, SUPER, or the vendor's name. Shared privilege is an exception unless the policy allows a named break-glass account, with a log of each use.
Evidence you keep. The role export, the signed review, and a cross-reference of items marked for removal against the later status. If the review is missing, keep the email that says it cannot be found. Absence is evidence too, when you record how you looked.
What an exception looks like. The quarter review is unsigned. Or it is signed, and NIGHT01 is still an admin on the reversal screen, with last login inside the night window. Or a leaver from the March list is still in the admin group. That last one links this test to WP-A12. Cross-reference the papers. Do not write the same finding twice and hope the marker thinks it is two tests.
Same day. Ask for the role export and last quarter's review pack. If the pack is a slide with no names, say so. Then ask for the system list anyway. The list is the population. The slide is a claim.
Change management on a live system
What the control is. This is an ITGC. No change reaches the live core, the switch, or the agent app without a request, a test, an approval by someone other than the developer, and a record of who moved it live. Emergency changes are allowed. They are not invisible. They need a ticket the same day or the next working day, and a look back after the rush.
Objective. To determine whether changes to the live core banking system in the period had a request, a test, and a separate approval before they went live.
Kind of test. Compliance. You compare what moved with what the change policy requires. Part 1's Thursday change to the night reversal limit is the picture to keep. The criterion is the bank's own change policy, not your taste about limits.
Population. Every production deployment in the period, taken from the deployment log or the change tool, not from memory. If developers can edit production without a deployment log, the population is incomplete, and that incompleteness is itself a result you record.
Sample. Test all emergency changes and all changes that touch payments, fees, limits, or access. Sample ordinary changes on top. A risk scored high in Part 2 means you do not sample only the calm releases and skip the night ones.
Steps.
Take the deployment list: change id, date, system, short description, and the name of the person who moved it.
Match each id to a ticket. XLOOKUP is enough. A deployment with no ticket is an exception. Write it down before you listen to the story.
On the ticket, check three things. There is a test note. The approver is a different person from the developer. The approval time is before the live time. An approval after the release is a story, not a control.
Open the tickets around the reversal program and the fee table. Look for the word temporary. A temporary limit change with no later ticket to put it back is an exception.
Ask whether anyone can log on to production and edit code or parameters with no ticket. If the answer is yes, try to see that path yourself, with permission, in a way that does not change anything. Observation plus the user list beats a denial in a meeting.
Evidence you keep. The deployment extract, the tickets for the sample, and screenshots or exports of the approval field. For the reversal limit, keep the parameter value you saw and the ticket that should explain it. Write the date you looked.
What an exception looks like. Release R-184 moved the reversal limit from UGX 2,000,000 to UGX 20,000,000 on a Thursday. The ticket says temporary. There is no second ticket. The developer and the approver are the same user. Any one of those facts is an exception. All three together are a control that did not operate.
Same day. Get the deployment list for one month first. A month is enough to learn the shape. Then widen to the period in the plan. Do not start with a six-month export you cannot open.
Backup and a tested restore
What the control is. This is an ITGC. The core is backed up on the schedule the policy names, failures are followed up, and someone tests a restore often enough that the bank knows the backup can be read. Part 1 was blunt about this. A recovery plan in a drawer, untested for two years, is a hope. Availability needs a restore you can point at.
Objective. To determine whether scheduled backups of the core banking system completed in the period, and whether a restore was tested in the quarter onto a system that was not serving customers.
Kind of test. Compliance. You are testing whether the operations control ran. A restore test that you observe, or a restore log that shows data opening, is stronger than a policy. You are not staging a disaster.
Population. Every scheduled backup job for the core in the period, and every restore test record in the quarter. If the policy says daily, the population is each day, not "the backup process" as a vague noun.
Sample. Review the job calendar for failed and missed jobs across the whole period. Those are the risky items, so take all of them. Then take one successful job from each week, so you are not only looking at bad news. Take the latest restore test in full. If there is no restore test, you do not invent a sample of backups and call the objective met.
Steps.
Read the backup standard once. Note the system, the frequency, and how often a restore must be tested. That is your criterion.
Export the job history: job name, start, finish, success or fail. Sort by date. Mark days with no job at all.
For each failure, find the ticket or the operator note. A failed row with no follow-up is an exception.
Ask for the latest restore test: date, who ran it, which backup was used, what was checked, and where it was restored. "We checked that the file exists" is not a restore.
Confirm the restore did not land on the live customer system. A test that overwrites live data is not a clever success. It is a different incident.
Evidence you keep. The job history export, the failure tickets, and the restore write-up. If the write-up has no date, say so. Keep the policy page that states the frequency, with the clause marked.
What an exception looks like. The job failed on 3 March and 4 March and nobody opened a ticket. Or the last restore test was fourteen months ago. Or the test report says the backup file was present and never says the data opened. The existence of a file is not a restore. You want evidence that a copy came back as a readable core, away from customers.
Same day. Ask operations for last month's job history and the last restore report. You can finish a first paper from those two documents in one afternoon. If they only send the policy, you are not done. You have the criterion. You do not yet have the condition.
Maker-checker on a reversal or a payment
What the control is. This is an application control. A reversal, or a payment above the floor, needs two people. The maker enters it. The checker approves it. The system should refuse to continue when the two ids are the same. Part 2 built the engagement around agency reversals for this reason. Customer balances move, and a reversal can hide a fraud. The score was 20 out of 25.
Objective. To determine whether agency banking reversals in the period were approved by someone other than the initiator, and logged to a named user.
Kind of test. Compliance. You are testing the operation of the approval control. The amounts matter later, as a substantive test, if this control fails or if the risk stays high. Do not start with the amounts and forget the ids.
Population. Every agency reversal in the period under review, from the reversal log, not from a spreadsheet a team rebuilt for you. In the Part 2 example the period is 1 April 2026 to 30 September 2026. Out of scope stays out: branch teller cash, ATM claims, loan recoveries, and the marketing website, unless a reversal in your set touches them.
Sample. Take every reversal at or above the materiality line of UGX 1,000,000. Then take a further sample of smaller reversals, because a theft can be split under the line. On a first day, if the log is huge, start with one month so you learn the columns, then expand to the period. Write that you started with a month, and do not pretend the month was the whole population.
Steps.
Export reversal id, date, time, amount, customer account, maker id, checker id, and channel. Save the original.
Add a column that compares maker and checker. Filter to rows where they match, and rows where the checker is blank.
Filter for shared ids. If maker and checker are both NIGHT01, you do not have two people. You have one password.
For a small set of passed rows, open the source screen with an officer and confirm the id on the log is the id on the screen. You are checking that the log is not a story written beside the real system.
If maker and checker differ, check they are not the same human with two ids. HR can show you that. A person who approves their own work under a second name is an exception.
Evidence you keep. The export, the filter results, the count of exceptions, and the policy clause that requires two people. Keep two or three passed items as well, so the paper shows you did not only hunt for bad rows.
What an exception looks like. Reversal AG-18820 for UGX 750,000 was made and approved by NIGHT01 at 01:14. Or the checker field is empty and the item still posted. Or the maker is J. Okello and the checker is J. Okello on a second user id created the same week. A busy night is an explanation. It is not a pass.
Same day. Ask for one month of the reversal log and the clause in the reversal procedure. You can run the same-id filter before lunch if the extract is clean. Write the count on the paper even if it is zero. A zero with no file name is not a zero.
Transaction limit
What the control is. This is an application control, and it acts at input. Each role has a ceiling. A teller may post up to a set amount. Above that amount the system blocks the payment or forces a second approval. For this illustration, the teller ceiling is UGX 5,000,000. Your case pack may use another figure. Use the figure in the parameter, not the figure a supervisor remembers.
Objective. To determine whether payments above the teller limit were blocked or sent for a second approval before they posted.
Kind of test. Compliance for the control itself. You also scan for posted items above the limit, which is a substantive check on the outcome. Do the scan. A limit that usually works, while three large payments posted with one id, did not work on those three.
Population. All payments initiated by teller roles in the month you are testing. The relevant cut of that population is every payment above the limit. Small payments do not exercise this control. Do not spend the day on them and call the limit tested.
Sample. Test 100 percent of payments above the limit. That is not a thin sample. It is the right population for this control. If that set is enormous, test all items above a higher line you set in advance, and a sample between the teller limit and that line. Set the line before you see which rows look odd.
Steps.
Read the limit from the parameter screen or the tariff table. Write the role, the amount, and the date you read it. Keep a screenshot of the parameter if you are allowed to.
Export payments: id, date, amount, initiator role, initiator id, approver id, status.
Filter to amounts greater than the limit. In the teaching case, greater than UGX 5,000,000.
On those rows, check for a second approver who is allowed to approve that size. A blank approver, or the same id, is an exception.
Do not try to break the limit on the live system to see what happens. If the bank gives you a test environment, one supervised attempt is enough, and it belongs in the paper as an observation. Production is not your lab.
Evidence you keep. The parameter extract, the filtered payment list, and the rows that posted above the limit. Note the export date and the person who ran it.
What an exception looks like. Payment P-44019 for UGX 8,500,000 was initiated by a teller id and posted with no second approver. Or the parameter says UGX 5,000,000, but the system posted a larger amount because a Thursday change raised the ceiling with no ticket. That second picture joins this test to change management. Cross-reference the papers.
Same day. One parameter screenshot and one month of payments above the line are enough to draft the paper. You can widen the period tomorrow. Record today what you actually held in your hands.
Fee or tax calculation
What the control is. This is an application processing control. The system should charge the fee in the tariff, and it should charge tax on that fee at the rate the rule states. In the teaching case the tax is 15 percent of the fee. Part 1 had the plain integrity failure: the tariff says 2,500 shillings, and the customer is charged 5,000. You do not settle that by arguing with the teller. You recompute.
Objective. To determine whether transfer fees in the month, and the tax on those fees, agree with the tariff and with tax at 15 percent of the fee.
Kind of test. Substantive. You reperform the calculation. This is not a test of whether someone signed a fee policy. A signed policy can sit beside a wrong formula. Recalculation is the point.
Population. Every fee-bearing transfer in the month, from the core, with the fee and the tax the system posted. The tariff table for that month is part of the evidence, not part of the population.
Sample. Recompute a sample of ordinary transfers, and recompute every transfer where the fee is large enough to matter. Use the Part 2 instinct. A single fee error at or above UGX 1,000,000 is material. A pattern of smaller errors that adds up to that amount is material too. On day one, recompute 30 ordinary rows plus every row above UGX 600,000 in fee, so performance materiality is already in the work. If the month has few rows, recompute all of them.
Steps. The Excel commands are written out in the tools section, so you can follow them key by key. The logic, before you touch a menu, is this.
Put the tariff beside you: product, fee, and the date the rate started. A tariff that changed during the month needs two rates. Do not use today's rate on last month's rows.
For each sampled row, look up the tariff fee. Compare it with the fee posted. The Part 1 case fails here when 5,000 was posted and the tariff says 2,500.
Compute expected tax as fee times 0.15. Compare it with the tax posted. Allow 1 shilling for rounding. A bigger gap is a flag.
Add the flags. Do not average them away. Ten small overcharges can cross the materiality line you set in planning.
If the formula in the system is wrong, note whether change control covered the last edit to that formula. A bad formula with no ticket is two tests pointing at one cause. You still write both results.
Evidence you keep. The tariff extract, the transaction extract, and the workbook where you recomputed. The workbook is evidence only if you still have the original export untouched. Keep both files.
What an exception looks like. Transfer T-90211 shows a posted fee of 5,000 and a tariff fee of 2,500. Or the fee matches, and the tax posted is 0 when 15 percent of the fee is 375. Or the fee matches the tariff and the tax matches 0.15, on the rows you picked, which is a pass on this sample. Write the pass with the same care as the fail. A blank conclusion is not a pass.
Same day. Build the workbook on 30 rows today. The SUMIFS total and the flag column should be working before you add more rows. A small correct workbook beats a large file you cannot explain.
Daily reconciliation of agent float or suspense
What the control is. This is an output control. After the day's agent deposits and withdrawals post, finance compares two records that should agree. One pair is the agent's float and the bank's record of that float. Float is the balance the agent is trusted to hold. Items that do not match sit in a suspense account until someone clears them. The control is that the comparison happens each business day, by someone other than the only person who can post adjustments, and that breaks are cleared rather than rolled forward in silence.
Objective. To determine whether agent float and the related suspense account were reconciled each business day in the month, and whether breaks above the agreed line were investigated.
Kind of test. Compliance on whether the reconciliation operated. Add a substantive piece: re-add a sample of packs yourself. An output control that exists only as a blank template did not operate.
Population. One reconciliation pack for each business day in the month, for the agent channel in scope. If the bank has many agents, the population can be the daily summary pack plus the exception list, not every agent's private notebook. Define that before you start, or the population will quietly become the packs they were happy to show you.
Sample. A month has about 20 to 22 business days. Test all of the daily packs if you can carry them. If you cannot, test every day where the break was at or above UGX 600,000, every day after a public holiday, and a handful of days that looked clean. Clean days stop you from reviewing only disasters. Missing days are exceptions, not days you skip.
Steps.
List the business days. Tick a pack present or absent. An absent day is an exception even if the next day is said to include both.
Check preparer and reviewer. They should be different people. The reviewer should sign after the preparer, on the same day or the next morning. A pack signed by one person twice is not a review.
Re-add the two sides of two or three packs. If the agent's device says one float and the core says another, the difference should equal the suspense item, or an explained timing item.
Follow one suspense item that is older than a day. Who owns it? Did it clear? A balance that rolls all month with no owner is an exception.
Tie one item back to a real deposit, using the reference from your walkthrough if you have one. You want to know the pack is about this channel, not about a different product with a similar name.
Evidence you keep. The list of days, the packs you reviewed, your re-addition, and the suspense report. Name the files. "Saw the recon" is not a file name.
What an exception looks like. There is no pack for 11 March. Or the 11 March pack shows a break of UGX 4,200,000 on agent float, prepared and reviewed by the same officer, and the same break is still open on 18 March. Or suspense holds an item labelled "agent difference" with no agent id. Part 1's agent with two phones, saying the float does not match, is the human picture. The pack is how you turn that picture into a test.
Same day. Ask for last week's five packs, not the whole year. Re-add one pack fully. If you can re-add one, you understand the test. Then scale it.
Understanding Audit Evidence
Audit evidence is the information you use to reach a conclusion. It is not the conclusion itself. "Access is weak" is a conclusion. The March extract, the two open accounts, and the policy clause are evidence. Part 1 put the difference in one small pair. An opinion says access seems okay. Evidence says which leavers you tested, what the rule was, and where the paper is.
Evidence has to earn its place. Three words are enough.
Relevant. It answers this objective, not a neighbouring worry. A photo of the server room does not answer a fee test.
Reliable. A careful outsider would trust the source, and could see that you did not alter it.
Sufficient. There is enough of it for the risk. One reversal, on a risk scored at 20, is not sufficient. All 25 leavers, on a list of 25, is sufficient for that compliance test.
The CISA Review Manual and ITAF expect you to judge evidence, not to collect souvenirs. A thick file can still be weak. Ten screenshots of the same policy page are one piece of evidence, repeated.
The reliability hierarchy
Some evidence is stronger than other evidence. Use this order when you decide what to rely on. It runs from stronger to weaker. You may still use the weaker kinds. You may not stop there when the risk is high.
What you do yourself. You recompute the fee. You match the leaver list. You watch a report run and take the file from that run. This is the top of the pile for most IS tests, because nobody stood between you and the result.
A record from outside the team that runs the process. A customer message that shows a different amount, a vendor notice, or a regulator return. Outsiders can be wrong. They are still harder for the auditee to edit quietly.
A system record you extracted, under a control environment you have some reason to trust. The core user extract, taken while you were there, with general controls that are not obviously broken.
A document prepared by the people who run the process, and emailed to you later. It can be true. It can also be the clean version. Use it to aim the test. Tie it back to the system before you rely on it.
A spoken answer, with nothing saved. Inquiry is the bottom of the hierarchy. It is useful. It tells you where the button is. On its own it does not support a conclusion. "The supervisor said reversals are always checked" is a lead. It is not a working paper.
Move up the list when Part 2 says the risk is high. Reversals at 20 do not rest on inquiry. Fees that can drift do not rest on a tariff poster. You recompute. You also write down the weakness of what you were given. If the only user list is an undated spreadsheet from a desktop, say that. Your conclusion should sound as strong as the evidence, and no stronger.
Reliability is not the same as friendly. The officer who stays late to help you can still hand you a filtered export. Thank them. Then check the row count against the system total. A helpful person and a complete file are different things.
The Two Kinds of Tests
Part 2 gave you the risk model in words. Audit risk is the risk of a wrong clean opinion. It depends on inherent risk, control risk, and detection risk. Fieldwork is where you deal with detection risk. You deal with it by choosing the right kind of test, and by doing enough of it.
There are two kinds.
A compliance test, also called a test of controls, asks whether the control operated in the period. Did leavers lose access in one working day? Did a second person approve the reversal? Did the backup job run, and was a restore tested? The answer is about the control, not about the total shillings.
A substantive test asks whether the data or the amount is right. You recompute the fee. You re-add the reconciliation. You scan for payments that posted above the limit. You are looking at the outcome. Substantive testing does not require the control to be pretty. It asks what landed in the ledger.
Tie the choice to the score you already wrote. When inherent risk and control risk are both high, you cannot keep audit risk low unless detection risk comes down. That means more testing, and better testing. For agency reversals, scored at 20, a short chat is not enough. You do the compliance test on maker and checker. If that test passes on a strong sample, you can lean on the control a little, and your substantive sample on the amounts can be tighter. If the compliance test fails, you stop leaning. You test more amounts. You do not pass the process because the policy reads well.
If general controls are weak, be slow to trust a compliance test that uses a report from that system. A maker-checker test on a spreadsheet the team typed is not the same as a maker-checker test on a log you extracted. Weak ITGCs push you toward substantive tests, and toward evidence you build yourself.
Students often do only one kind, and they pick the comfortable one. They collect signatures and never recompute. Or they recompute thirty fees and never ask who can change the fee table. The program should show both, in proportion to the risk. A low score, such as the branch poster at 2, gets neither. You recorded it in planning. You do not build a sample for it now.
Compliance tells you whether the net was used. Substantive tells you whether anything tore through it.
Sampling
A population is the full set of items the conclusion is about. A sample is the subset you actually test. The conclusion still speaks about the population. That is why a sample you cannot explain is dangerous. You will sound as if you tested the bank, when you tested the rows someone put on a memory stick.
You do not always sample. If the population is 25 leavers, test 25. If the control only matters above a limit, test the items above the limit. Sampling is what you do when the full set is too large for the risk, not a ritual you perform on every list.
How big is big enough? Part 2 already answered in words. High inherent risk and high control risk mean a lower detection risk, which means a deeper sample. The teaching picture was numeric. With inherent risk 0.8 and control risk 0.9, detection risk had to fall to about 0.14 if you wanted audit risk near 0.1. You will not compute a statistics formula on the capstone unless the brief tells you to. You will write the decision. Reversals are scored at 20, so the sample includes every item at or above UGX 1,000,000, plus 40 smaller items chosen without hunting for easy rows. That sentence is better than a magic number with no reason.
Three ways of choosing are enough for this course.
Test all when the population is small, or when every item above a line matters.
Random when any item in a large list could carry the failure, and you do not want to pick only the ones that look odd. State the method. Every 50th row after a random start is a method. "I clicked some" is not.
Judgmental when you have a reason that belongs to the risk. All emergency changes. All days after a holiday. All shared ids. Judgmental is allowed. Unexplained is not. Write the reason before you see the result, so you are not fitting the sample to the mistake.
An exception is a sampled item that does not meet the criterion. One exception does not always mean the whole population failed. Zero exceptions do not always mean the whole population is clean. What you may say depends on how you sampled, and on how serious the miss is. A single shared admin id can be enough to call the control ineffective, because of its nature. A single fee that is 1 shilling out, in a random sample, is a different conversation. Materiality from the plan is how you tell those apart. Do not invent a new threshold after you see the row.
The four things you must document
Whatever method you use, the working paper has to let a reviewer rebuild your thinking. Document these four things every time. If one is missing, the sample is not finished.
The population. What it is, which system, which period, and how many items. "Reversals" is not enough. "Agency reversals on the core from 1 April 2026 to 30 September 2026, 6,420 rows in reversal_log.csv" is enough.
The size and why. How many you tested, and the link to the risk score, the materiality line, or the fact that the population was small enough to test in full.
The method and the items. All, random, or judgmental, plus the list of items or the rule that rebuilds the list. A reviewer who cannot get back to the same rows cannot review you.
The exceptions and the conclusion about the population. Count them. Describe them. Say what you now believe about the population, not only about your favourite rows. If you found two open leavers out of 25, you do not conclude that access removal mostly works and file the paper as a pass.
Keep the original population file. If you delete the rows you did not like, you no longer have a sample. You have a scrap.
CAATs
CAAT means a computer-assisted audit technique. It is any use of a computer to do audit work that you would otherwise do slowly by hand: match two files, recompute a column, group a population, or flag duplicates. The computer does not conclude. You do. The tool makes the test repeatable.
For the capstone, Excel is enough if you can explain every formula you rely on. You will also hear the names IDEA, ACL / Diligent Analytics, SQL, and Benford. Those names are part of the wider toolkit. You do not need them to finish this part of the course. Learn Excel properly first. A named tool you cannot explain is worse than a simple formula you can rerun in front of the class.
Excel, in real steps
Use a copy. The original export stays untouched, so you can always get back to what the system gave you. This walk-through uses a fee file. Column A is the product code. Column B is the product name. Column C is the fee the system posted. Column D is the tax the system posted. The tariff sheet is called Tariff, with the product code in column A and the tariff fee in column B.
Follow the menus as they are named in current Excel. The commands below are Filter, PivotTable, XLOOKUP, and SUMIFS.
1. Save the export as fees_march.csv. Copy it. Name the copy fees_march_work.xlsx. Do not type over the original.
2. Open fees_march_work.xlsx. Click any cell in the data. On the Data tab, click Filter. Filter arrows appear on the headings.
3. Open the Filter arrow on Fee. Untick (Blanks) so you can see whether any transfers posted with no fee. Note the count. Clear the filter when you have written the count down.
4. Click any cell in the data. On the Insert tab, click PivotTable. Put it on a new sheet. Drag Product into Rows. Drag Fee into Values. Drag Tax charged into Values. Both should read Sum. If Excel says Count, click the value and change it to Sum.
5. Back on the data sheet, look up the tariff fee with XLOOKUP. In F2 type:
=XLOOKUP(A2,Tariff!A:A,Tariff!B:B,"No tariff")
6. Total the fees for one product with SUMIFS. In a cell to the side, type:
=SUMIFS(C:C,B:B,"Transfer")
7. Expected tax is the tariff fee times 0.15. In G2 type:
=F2*0.15
8. Flag a difference greater than 1 shilling. In H2 type:
=IF(ABS(D2-G2)>1,"Flag","OK")
9. Fill F2, G2, and H2 down to the last row. Turn Filter on again. Filter column H to Flag. Those rows are the tax exceptions. Save the workbook. Write the file name on the working paper.Read the flag before you celebrate it. XLOOKUP brings the tariff fee. If it returns "No tariff", you have a product the tariff does not know. That is an exception of a different kind. SUMIFS adds the posted fees for the product named Transfer. Use it to compare the detail rows with the pivot total. If those two totals disagree, you have filtered a different set, or you have included a blank. Do not ignore a disagreement. It means your population moved while you were working. The tax flag compares tax posted in column D with expected tax in column G. One shilling is rounding. More than that is a flag you open, not a flag you hide with another filter.
Two habits keep Excel honest. First, do not paste values over the formulas and throw the formulas away. The formula is the test. Second, when you sort, sort the whole table. A sort on one column misaligns the fee and the customer. That false exception will embarrass you in the review, and it should.
A small pandas check for leavers
Pandas is a Python library for tables. You do not need it if Excel is working. It is useful when the leaver list and the user list are large, and you want the match to be obvious. This script keeps only the people who appear on both files and whose account is still active. Those are the leavers who should have been disabled.
import pandas as pd
leavers = pd.read_csv("leavers_march.csv")
active = pd.read_csv("active_users.csv")
matches = leavers.merge(active, on="staff_id", how="inner")
should_be_disabled = matches.loc[matches["status"] == "Active"].copy()
should_be_disabled.to_csv("wp_a12_exceptions.csv", index=False)
print(should_be_disabled.shape[0])The inner merge is the part to understand. how="inner" keeps only staff ids that are on both lists. A leaver with no user row drops out. A user who did not leave drops out. What remains are leavers who are still on the user file. The next line keeps status Active. Those rows should have been disabled. The csv it writes is an evidence file. It is not the conclusion until you check that staff_id means the same thing on both extracts. Run it. Then open the csv and read the names. A script you have not opened is a rumour.
Continuous auditing, briefly
Continuous auditing means tests that run often, sometimes every day, instead of once in a yearly visit. A daily report of leavers who are still active is continuous auditing. A daily flag of reversals where maker equals checker is the same idea. It helps management see breaks while they are small. It does not replace your working paper. If you rely on that daily report, you still record which days you took, what the report contained, and what you concluded. A dashboard with no owner and no retained result is a screen. It is not assurance.
Working Papers
A working paper is the record of one test. It shows the objective, the work, the evidence, and the conclusion. Part 1 called working papers the audit's memory. That memory has to be good enough for a reviewer who was not in the room, and for you in three weeks, when the exit meeting asks a sharp question.
Index the papers. WP-A12 means working paper A12. The letter can mark the process. The number marks the test. Use one reference for one test. Do not keep notes in a personal notebook that never joins the file. The capstone file is the notebook.
A paper that another student can review has these fields. Leave none of them blank. If the honest entry is "not received", write that.
Reference and title. The index code, and a name a stranger understands.
Objective. The sentence that begins "To determine whether".
Criteria. What should be. The policy clause, or the rule you are applying.
Period. The dates the conclusion covers.
Population. The full set, the source, and the count.
Sample. The size, the method, and the four sampling facts from the previous section.
Work done. The steps, in the order you did them, including the formula or the match.
Evidence. File names, who provided them, the date, and the rows you rely on.
Result. The exceptions, or a clear zero.
Conclusion. Effective, or not effective, for this objective, this population, and this period. Say what you cannot conclude.
Preparer and reviewer. Two people. The reviewer is not a signature added in the corridor with the paper unread.
Cross-reference. The program step, and any other paper that shares this evidence.
Write the conclusion in the same words you would say to the audit committee, only shorter. "We matched all 25 March leavers to the core extract. Two accounts were still active. For this population the control did not operate." That is a conclusion. "Issues were noted" is a fog.
Part 1 showed a clean version of this test, so you could see what evidence looks like when the control holds: 25 leavers, access removed within one working day, recorded on working paper A-12. The paper below is the same test, with the reference written as WP-A12, and it shows the version you also need to know how to write. Two accounts are still active. The illustrations are not a claim that a real March looked like this. They are a filled example you can imitate in structure.
Filled example: WP-A12
Field | Entry |
|---|---|
Reference | WP-A12 |
Title | Leaver access removal, core banking, March |
Objective | To determine whether access for staff who left in March was removed within one working day. |
Criteria | Access removed within one working day of the last working day. Source: BBC Bank access standard, teaching clause used in class. |
Period | 1 March 2026 to 31 March 2026. |
Population | 25 leavers in March, from hr_leavers_march.csv, matched to the core user extract. |
Sample | All 25. The population is small and the risk affects customer data, so the whole population was tested. |
Work done | Matched all 25 leavers to core_users_extract.csv on staff id. Compared last working day with status and disabled date. Counted working days, so a Friday leaver may be disabled on Monday. |
Evidence | hr_leavers_march.csv (25 rows) and core_users_extract.csv, run on 31 March 2026. Exception detail is on rows 14 and 22 of the match file wp_a12_exceptions.csv. |
Result | 2 exceptions: accounts still active. Staff 1044, last day 12 March, status Active. Staff 1182, last day 6 March, disabled only on 20 March, which is outside one working day. |
Conclusion | Not effective. For this population of 25, access was not removed within one working day in two cases. No wider conclusion is drawn about other months. |
Preparer | Your name |
Reviewer | Reviewer |
Read that table as a form, not as a story. A reviewer can see the objective, the sample, the evidence, and the conclusion without asking you a single question. That is the standard. If your capstone paper cannot be read that way, it is not finished.

Walkthroughs
A walkthrough is not a test. Part 2 was firm about this, and fieldwork is where students forget it. A walkthrough follows one real transaction from the start to the finish, with the person who does the work. You are learning the path. You are not selecting a sample. You are not arguing. If you correct people during the walk, they will show you the official path and hide the path they use on a busy Saturday.
Here is one agent-deposit story. It is an illustration. Use it to practise looking, not as the bank's script.
On a Tuesday you sit with Amina, an agent. A customer hands her UGX 50,000 in cash for a deposit. She counts the notes twice. She opens the agent app, enters the phone number, and waits until the customer's name appears. She asks the customer to confirm the name. Then she enters 50,000 and submits. The phone shows a success screen and a reference, AG-20441. The customer's phone beeps. You write the time, the reference, the screen name, and the fact that she counted the cash before she submitted. You do not ask her to change anything. You thank her and you leave the queue alone.
Later the same day you ask operations to show you AG-20441. You see the message hit the switch. You see the core credit the customer and reduce the agent float. You ask, without accusing anyone, who could reverse this item tonight, and whether that person could also approve the reversal. You ask which report tomorrow morning should show this deposit if the float is reconciled. You write the names of the screens. You stop.
That narrative is a walkthrough. It becomes fieldwork only when you pick a population. The deposit you saw is one item you now understand. It is not a sample of one on which you conclude that all agent deposits are sound. Hang later tests on what you saw.
If Amina could have reversed her own deposit, you have a lead for the maker-checker test.
If the amount had been far above a normal agent deposit, you have a lead for the limit test.
If the fee on the receipt disagreed with the tariff, you have a lead for recomputation.
If tomorrow's pack does not contain AG-20441, you have a lead for the output reconciliation.
Write the walkthrough on one page: who you sat with, the item you followed, the screens, and the controls you expected and did not see. Then open a different working paper for the test. Mixing the two is how a story gets dressed up as assurance.
Putting It Together
Fieldwork is a loop, not a pile of files. You pick the next test from the program. The program came from the plan. The plan came from the risk. If you cannot point to that chain, you are wandering, and Part 2 already showed you what wandering costs. The room goes quiet when someone asks what conclusion the sample was meant to support.
For each test, decide whether it is compliance or substantive. Choose the sample and justify it with the score, the materiality line, or the small size of the population. Run it with a tool you can explain. Excel is enough. Judge the evidence with the hierarchy in mind. Inquiry starts you. It does not finish you. Then write the working paper before you start the next test. A paper written a week later is a reconstruction. Reconstructions smooth out the awkward bits. The awkward bits are often the finding.
The conclusion on the paper is one of two things. The control was effective for this objective, on this sample, in this period. Or there is an exception, and you say what it was. You do not invent a third result called "mostly fine" and hope the review is rushed. Effective, or an exception. If you found an exception, you do not talk yourself out of it in the corridor. You record it. Severity, the report, and the argument about what management should do come in the next part of this series. Your job here is to make that next part possible.
Then you go back to the program and pick the next test. Access before you lean on the fee formula. Maker-checker on the reversals you scored at 20. A recompute where the money is calculated. A reconciliation where the day is supposed to close. When the program steps are all backed by papers, fieldwork is done. Not when you are tired. When the file can stand without you.
Read the figure as a loop. Start at the first box. Move to the right. The arrow brings you back for the next test. Do not skip the paper and jump straight to another extract.

Key Terms
Fieldwork. Phase 2. You carry out the tests and record the evidence. It is detective work, not a tour, and not a repair job.
Audit program. The document that says how you will test. The plan said what and why. The program says how.
IT general controls. Controls around access, change, and operations that many systems share. Test them before you lean on one application.
Application controls. Controls inside one process: input, processing, and output.
Maker-checker. The person who starts an item is not the person who approves it.
Limit. A ceiling the system enforces, by blocking the item or by requiring a stronger approver.
Compliance test. A test of whether the control operated as required.
Substantive test. A test of whether the amount or the data is right, often by recomputation.
Population. The full set of items your conclusion is about.
Sample. The items you actually tested. You must be able to say how you chose them.
Audit evidence. The information that supports a conclusion. It should be relevant, reliable, and sufficient.
Working paper. The record of one test: objective, work, evidence, and conclusion. If it is not written here, it did not happen.
CAAT. A computer-assisted audit technique. Excel, used with a formula you can explain, is the CAAT you will use first.
Continuous auditing. Tests that run often, not only at the yearly visit. They still need a retained result.
Walkthrough. Following one real transaction so you understand the path. It is learning. It is not a sample.
Exception. An item that does not meet the criterion. Record it. Do not smooth it into a soft pass.
References
ISACA (2024). CISA Review Manual, 28th edition. Schaumburg, IL: ISACA.
ISACA (2020). IT Audit Framework (ITAF), 4th edition. Schaumburg, IL: ISACA.
Hall, J. A. (2015). Information Technology Auditing, 4th edition. Boston: Cengage Learning.
Davis, C., Schiller, M. and Wheeler, K. (2011). IT Auditing: Using Controls to Protect Information Assets, 2nd edition. New York: McGraw-Hill.
Coderre, D. (2009). Computer-Aided Fraud Prevention and Detection: A Step-by-Step Guide. Hoboken: Wiley.
Your Reflection and Peer Review
This section is a required learning activity. Fieldwork is practical, so your reflection should show that you can picture yourself doing it.
Part A: Write your own reflection (about 300 to 400 words)
In the comments below, post a reflection answering the following.
In your own words, explain the difference between a compliance test and a substantive test, with one example of each.
Explain why evidence you generate yourself is more reliable than a manager telling you something. Give an example.
Describe one CAAT you could perform in Excel, step by step, on any dataset you can imagine. Name the Excel feature you would use.
List your three key takeaways from this article and why each matters.
Describe one tool or technique in this article that you want to learn more about, and why.
Part B: Critique at least five peers' reflections
Read your classmates' reflections and respond to at least five. For each, write two to four sentences that do the following.
Name one thing they explained well and say specifically why.
Identify one gap or error, especially if they confused compliance and substantive testing, or described a CAAT that would not actually work, and suggest how to fix it.
Ask one real question that pushes their understanding further.
What a good critique looks like
Weak: "Great example, I would do the same." Strong: "Your Excel CAAT for finding duplicate payments is a good idea, and using COUNTIF to flag repeated reference numbers is exactly right. But you called it a compliance test. Checking whether a payment was made twice is really a substantive test, because you are testing the numbers, not whether a control operated. How would you turn it into a compliance test of the approval control instead?"
Ground rules
Critique the idea, not the person. Be specific and constructive. The goal is for the whole class to become confident with evidence, testing and tools.