The engineering moved faster than the commercial thinking.
A critique of 61 graduation projects from the 2025/26 professional training programmes, the catalogue they sit in, and the freelancing outcomes reported alongside them. Every figure here can be opened down to the slide it came from.
Six branches did not submit
Assiut, Banha, Mansoura, New Capital, New Valley and Port Said are absent from every figure in this report. They represent 23% of ITP offerings and 15% of PTP. Read every headline as 16 of 22 branches reporting.
How to use this
- Click anything. Every dot, bar, cell and table row opens the full record behind it — scores with their justifications, comparables, cost assumptions, and the slide reference it was extracted from.
- Press P for presenter mode: one exhibit fills the screen, ← → step through 18 of them, Esc exits. Records still open while presenting, so a question from the room gets answered by clicking the dot.
- Press Ctrl+K to jump to any project, branch, criterion or competitor by name.
- Printing gives the narrative, the charts and the tables. The drill-down layers are screen-only — a printout is the summary, not the archive.
What we found
The headline is not the one we expected going in. We assumed this review would mostly document AI-washing — projects claiming agentic AI that turned out to be chatbots. It did not.
Twenty-one of 61 projects claim agentic AI. Twenty of them can back the claim with a described architecture: named agents, tool use, plan-act-observe loops. Students across the network can genuinely build this.
The weakness is on the other side of the ledger. 14 of 61 projects state how they would make money. 2 state how they would reach a customer. 35 are entering markets our research rated crowded or saturated, and only 9 address a gap incumbents have actually left open.
The nine measures
Every tile opens the exhibit behind it. Figures cover 16 of 22 branches.
In one line
The engineering has moved faster than the commercial thinking, and the gap is now the binding constraint on what these projects can become.
The agentic AI picture
We scored two things separately and deliberately: what was actually built, and what the problem genuinely required — judged from the problem alone, before looking at the build. Comparing them gives four quadrants.
Agentic depth built against agentic necessity
Every dot is one project; click it to open the full record. Cells lay their projects out as a small grid, so counts stay exact where projects share coordinates. Use the legend to isolate a quadrant.
Only 2 projects over-engineered. That is a good result, and it contradicts the common assumption that students bolt agents onto everything. The 7 under-powered projects are the ones worth acting on — real problems where the team reached for a chatbot when the problem justified more.
The under-powered seven
| Project | Branch | Needed | Built |
|---|---|---|---|
| DAR Platform | Damanhour | 3/5 | 1/4 |
| Hiring Stage | Damanhour | 3/5 | 1/4 |
| NetFix AI | Ismailia | 3/5 | 1/4 |
| NutriScan AI | Smart Village | 3/5 | 1/4 |
| Training Simulation | Smart Village | 3/5 | 1/4 |
| Mongez | Smart Village | 3/5 | 1/4 |
| MedAgents | Sohag | 3/5 | 1/4 |
These are teachable, not failures. Each already has a problem worth solving, with the validation done — which makes them ready-made briefs for the next intake.
How the quadrants are derived
C3 agentic depth is scored from the described architecture only — the word “agent” in a bullet is not evidence, a plan→act→observe loop with tool calls is. C4 agentic necessity is scored from the problem statement alone, before looking at what was built. The quadrant is then mechanical: over-engineered is C3 ≥ 2 with C4 ≤ 2; under-powered is C3 ≤ 1 with C4 ≥ 3.
A stack monoculture is forming
The concerning reading is not that RAG is popular. It is that RAG appears on problems that do not obviously need retrieval, which suggests one template being applied regardless of the problem.
What the portfolio is built from
Counted across the 45 projects that named a stack at all. Click any bar to list the projects behind it.
RAG appears in 24 projects across 13 branches and LLM in 14 across 8. Behind them sits the same supporting cast — Qdrant, Ollama, Semantic Kernel, Gemini — in near-identical combinations. Two Damanhour projects list identical AI stacks.
How these counts were derived
A term counts once per project where it appears as a whole word in the technologies or AI-components text as submitted. Substring matching on a word boundary is deliberate: the source cells concatenate tools, so Qdrant(RAG) and chatbot – Qdrant – RAG are both genuine RAG mentions that a delimiter split would miss. The companion curriculum document reports lower counts because it split on delimiters only — these figures are the more complete of the two.
Recommendation
Teach the choice, not just the tool. Ask every team to justify why their AI approach fits their problem, and to name the simpler approach they rejected. Include at least one project brief where the right answer is deterministic code.
The commercial gap
The same teams that can build a multi-agent system cannot say who pays, how much, or how the product reaches anyone. This is the largest single finding in the review.
Every project, by what it says about money
One cell per project. 14 of 61 state a revenue model; 2 also state a go-to-market; 6 give enough for unit economics to be assessed at all. Click any cell for the record.
Two projects in 61 described a route to a customer. That is not a student failing — it is not currently being asked of them. A one-page business canvas as a mandatory graduation artefact would close most of this gap, and it is a week of teaching against a weakness that affects 77% of projects.
Why these are not zeroes
A project with no stated revenue model scores NE on B1, not 1. NE means the source did not say; 1 means stated, and bad. Conflating them would punish branches for thin documentation rather than thin thinking, so NE is excluded from the merit average and counted on the evidence axis instead.
Nobody checked who else was building this
35 of 61 projects entered markets our research rated crowded or saturated. Only 9 addressed a gap incumbents had genuinely left open.
Crowdedness of the markets entered
Scored per project against named global leaders and Egyptian players, citing 426 sources across 180 comparables.
Five projects landed on top of a funded incumbent
| Project | Branch | Incumbent | Why it matters |
|---|---|---|---|
| Furnora | Damietta | Homzmart | $15M Series A, already operating in Damietta furniture |
| Elshamy Pharmacies | Smart Village | Chefaa | 1.5M+ monthly users; El Ezaby runs 185+ branches |
| Brixel | Minia | Kuadra | Egyptian AI construction platform, launched 13 July 2026 — days before submission |
| On-Demand Home Nurse | Smart Village | 7keema | billed as Egypt's first on-demand home nursing app |
| CLINIKA | Ismailia | AI4Docs.AI | Arabic ambient clinical scribe built in Cairo |
Two hours of searching at proposal stage would have redirected several of these toward genuine whitespace. Open any project to read the full comparable set, including pricing and where the incumbent falls short.
Recommendation
Require a prior-art check before a project is approved, not after it is built.
The expensive projects are the ones with no revenue model
Every project was modelled at a common 1,000 monthly active users so the figures stay comparable. These are order-of-magnitude estimates with stated assumptions, not budgets.
What the median is taken over
15 of the 61 projects describe no AI at all (C1 = 0) and therefore cost nothing to run. They are excluded from the median, which is taken across the 46 projects that do use AI — including them would drag it to $200 and describe a “median AI running cost” using projects that run no AI. The portfolio total below keeps all 61.
Every project by estimated monthly AI cost
Bars are coloured by whether the project states a revenue model. The pattern at the top of the chart is the finding: the most expensive projects are disproportionately the ones with no stated way to pay for themselves. Open any bar for its full assumption set.
The five above $1,000 per month
All five are driven by speech-to-text or multimodal processing on clinical and media workloads. Two carry an explicit caveat that 1,000 MAU is artificial for a B2B clinician tool — surfaced in the record rather than hidden.
Recommendation
Teach cost modelling alongside AI architecture. A team that can estimate its own inference bill will design differently, and will find the constraint before a customer does.
Reporting standards vary more than project quality
Evidence quality was scored separately from project merit precisely so branches would not be ranked on slide design. That separation turned out to matter a great deal.
Average evidence score by branch
Smart Village submitted the most projects (19) and documented them the least, averaging 1.86 against 3.40 across every other branch. 16 of its 19 could not be assessed on enough criteria to rank fairly.
This is not evidence that Smart Village's projects are weak. It is evidence that we cannot tell — which, for the branch running 51% of PTP Intake 46, is its own problem. Its game and data-management projects arrived as feature tag-clouds with no problem statement, customer or business model. Alexandria's entire deck yielded 36 characters of machine-readable text. Damanhour and Qena submitted structured canvases. All were responding to the same request.
Documentation bands
| Band | Projects | What it means |
|---|---|---|
| A — full | 8 | Scored on 12 or more of the 15 merit criteria. The only projects rankable with confidence. |
| B — partial | 30 | Enough to assess, not enough to rank against band A. |
| C — insufficient | 23 | Too little documented to apply most of the rubric. Not a judgement on the project. |
23 of 61 projects sit in band C. Compare within a band, not across one.
A caveat we are not dropping
Merit score correlates with documentation depth at +0.68. Two readings are available and both are probably partly true: teams that thought harder produced better projects and better documentation; and our score partly rewards documentation depth regardless of merit. That is why projects are banded, and why the ranking below covers only band A.
Ranked — the fully assessable projects
8 projects had enough documentation to score on 12 or more of the 15 merit criteria. These are the only projects that can be ranked with confidence.
Band A, ranked by merit
| # | Project | Branch | Merit | Criteria | Evidence | Quadrant |
|---|---|---|---|---|---|---|
| 1 | Arena — AI-Powered Gym Management System | Qena | 3.86 | 14 | 4.33 | aligned agentic |
| 4 | ShopBrain Analytics | Sohag | 3.58 | 12 | 3.67 | aligned agentic |
| 9 | DAR Platform | Damanhour | 3.43 | 14 | 4.33 | under powered |
| 11 | Smart Feedback | Aswan | 3.36 | 14 | 4.00 | aligned simple |
| 16 | Hiring Stage | Damanhour | 3.23 | 13 | 4.33 | under powered |
| 21 | TaskPilot | Aswan | 3.08 | 12 | 4.00 | aligned agentic |
| 33 | AURA AI | Ismailia | 2.83 | 12 | 3.67 | aligned agentic |
| 38 | Brixel | Minia | 2.79 | 14 | 3.67 | aligned agentic |
Merit is the mean of the scored merit criteria, excluding NE. “Criteria” is how many of the 15 could actually be applied — the number that makes the merit score trustworthy or not.
Arena — AI-Powered Gym Management System leads on a full 14 criteria — an AI gym management system whose deterministic clinical-safety layer validates AI-generated workout and nutrition plans against ACSM/AHA/ADA guidelines, a guard none of its three comparables implements.
The other 53 projects are not ranked here
They are all in the explorer below, with every score, justification and comparable intact. They are unranked because their documentation could not support a rank, which is a different statement from scoring badly.
Every project
All 61 projects, with every score, justification, comparable and cost assumption intact. Filter, sort, then open any row for the complete record.
The full portfolio
Sort by any column. Merit is only comparable within a documentation band — the “of 15” column tells you how much of the rubric could actually be applied.
Unattributed records are kept, not dropped
1 project carries no branch and 1 carries no programme; 38 of 61 carry no track, because most decks never printed one. These appear under an “unattributed” option in each filter rather than disappearing from the counts. Track recovery went from 4 to 23 during the image pass; the rest genuinely are not in the source.
By branch
16 branches submitted. Each row opens that branch's full record — its projects, its rollups, and its freelancing rows.
Reporting branches
Averages cover only that branch's own projects. Branches submitting one or two projects have averages that should not be read as a ranking.
The six that did not submit
These branches are absent from every figure in this report. Between them they run the offerings listed here, which is the size of the blind spot.
Why PTP cannot be compared across branches
PTP Intake 46 runs in only 11 of 22 branches, and 28 of its 55 tracks — 51% — sit at Smart Village. Any cross-branch PTP comparison is close to meaningless; use department and area instead.
Freelancing — reported, not yet reconciled
The freelancing figures branches reported alongside their projects. This is raw data with its quality problems left visible: the reconciliation against the system, and the job-type and platform analysis, are a separate piece of work that has not been done.
Read this before quoting any number here
These 119 rows are what branches stated. Every row is marked derived: N — nothing here was computed by us. 12 rows fail an internal consistency check and 3 income figures are statistical outliers, all listed below rather than quietly cleaned. Rows also differ in grain — some are per-track, some per-branch, some per-person — so they must not be summed across grains or compared against a per-track system figure.
Known data problems
Recorded during the merge and carried through unchanged. Resolving these is the first step of the reconciliation, not something to do silently afterwards — reconciling against known-bad numbers wastes the exercise.
Freelancing summary — every reported row
Grain is shown on every row. Where a track name could not be matched to the catalogue, the raw string as written on the slide is shown instead.
Named freelancers
38 individuals with a recorded project. Platform is missing for 11 of them. Open a row for the project idea, tools and client.
What still has to happen
Obtain the system export and confirm its grain and columns; resolve the arithmetic and percentage failures above; join on offering_id and produce a variance table of reported against system; and only then analyse job types and platforms. Zagazig's single “Total $35,857” must be split by programme first.
The programme catalogue
The cleaned reference the whole report joins against: 22 units and 219 offerings, keyed by offering_id (programme × department × track × branch).
Offerings
Both programmes in scope: PTP Intake 46, and ITP 25/26 with rounds 1 and 2 merged.
Source files
The 22 branch and department files this report was built from. Much of the content was locked inside slide images and was recovered by re-reading every referenced slide at 150 DPI.
Open gaps
What would have to be requested to close the record. Blocking gaps are the six missing branch reports.
Method, verification and limits
How the scores were produced, what was done to check them, and what this report cannot tell you. None of this is buried — the caveats are part of the finding.
The NE rule
NE is never a low score
NE means the source material did not say. A 1 means stated, and bad. Conflating them would punish branches for poor documentation rather than poor thinking, so NE is excluded from the merit average and counted on the evidence axis instead. Inference is not evidence: if a target market can only be reverse-engineered from a solution sentence, A2 is NE. This rule underpins the whole two-axis design.
The rubric
27 criterion identifiers, of which 23 carry a 1–5 score and 4 are structured fields rather than scores — C2 records the agentic claim as written, C5 derives the quadrant, C6 flags AI-washing, and C7 holds the cost model. Not every criterion applies to every project: the number actually applied ranges from 21 to 23, and a record distinguishes scored, NE and not applied as three different states.
Every criterion, with its anchors
Expand any criterion for the full scoring anchors as they were applied.
Verification
An independent agent re-scored a stratified sample of nine projects spanning the full documentation range, without sight of the original scores. Across 180 comparable score pairs: 73% exact agreement, 90% within one point, and 8 of 9 quadrants agreed. Quadrant agreement matters most, since the quadrant is this report's central claim — and the verifier independently reproduced TAGER as over-engineered, the counterintuitive call in the original pass.
The automated NE audit found zero violations: no criterion was scored 1 where the underlying source field was empty, which is the exact failure the rubric was written to prevent.
An inconsistency we found and did not hide
The blind re-score disagreed on whether a criterion was NE or scoreable in 16 of 180 pairs (9%), one-directionally: on thin projects the original pass scored E3 and D3 numerically where the verifier judged NE. Merit was recomputed with the disputed criteria removed. The mean shift was +0.00, the largest single move was +0.38, and no project changed band. The scores were therefore retained and the finding recorded here rather than silently corrected.
Limits
What this report cannot tell you
| Limit | What it means |
|---|---|
| Six branches did not submit | Assiut, Banha, Mansoura, New Capital, New Valley, Port Said — 23% of ITP offerings and 15% of PTP. |
| Merit correlates with documentation at +0.68 | Projects are banded and must be compared within a band. Stated openly rather than dropped. |
| 38 of 61 projects have no track | Most decks never printed one. Department-level analysis is therefore partial. |
| Prior art is research, not audit | 180 comparables across 426 sources, gathered per project. Search quality is recorded per project and shown in each record. |
| AI costs are order-of-magnitude | All modelled at a common 1,000 MAU so they stay comparable. Two carry an explicit caveat that 1,000 MAU is artificial for a B2B clinician tool. |
| The JETS department mapping is unconfirmed | Mapped to Java and Mobile Native at medium confidence. |
Every one of these is live. None is a reason to discard the findings; all are reasons to read them at the right resolution.