Table of Contents
Portfolio Assessment vs Standardized Tests: What Each Measures and What Each Misses
Portfolio-based learning and standardized tests measure genuinely different things. Here's what the research says about each approach, their equity implications, and what hybrid assessment looks like in practice.
The debate over standardized testing in American schools is almost always framed as a political one. That framing obscures something more useful: standardized tests and portfolio-based assessment are not competing methods for measuring the same thing. They are measuring genuinely different things. Understanding what each captures — and what each cannot capture — is more useful than choosing a side.
This distinction matters more now than it did a decade ago. Several large school systems, including the University of California system, shifted to test-optional policies during the pandemic. New York City has expanded portfolio-based assessment. Meanwhile, standardized testing remains the backbone of federal accountability through Every Student Succeeds Act requirements. Parents navigating decisions about schools, programs, and educational priorities deserve a clear-eyed account of what the research actually shows.
Key Takeaways
- Standardized tests reliably measure a specific set of academic skills — primarily reading comprehension, mathematical computation, and procedural knowledge — and predict future academic performance reasonably well within that domain.
- Portfolio assessment captures learning growth over time, creative and applied problem-solving, student voice, and the development of skills that standardized tests do not assess.
- Both approaches have documented equity problems, but they operate differently: standardized tests show persistent demographic score gaps; portfolio assessment shows persistent implementation quality gaps that disadvantage students in under-resourced schools.
- Neither approach alone gives a complete picture of student learning. The strongest evidence supports hybrid models that use each method for what it actually measures well.
- The political fight over testing is partly a proxy fight over which skills schools should prioritize — and that is a legitimate values disagreement, not a measurement dispute.
What Standardized Tests Actually Measure
The Case for Standardized Tests
Standardized tests were designed to solve a real problem: how do you compare student performance across different teachers, schools, and districts when grading standards vary? A B in one classroom might represent very different work than a B in another. A test that every student takes under identical conditions, scored by the same rubric, provides a form of comparability that teacher-assigned grades cannot.
The research on what standardized tests predict is reasonably robust. SAT and ACT scores correlate meaningfully with first-year college GPA — the correlation is typically around 0.35–0.50 in large studies, which is modest but statistically significant. State standardized tests in grades 3–8 predict later course grades and graduation rates at the population level. The tests do measure something real.
What do they measure? The most careful analyses suggest that standardized tests primarily capture:
- Declarative and procedural knowledge (facts, formulas, conventions)
- Reading comprehension of informational and literary texts under time pressure
- Mathematical computation and algebraic reasoning
- Familiarity with academic language and test-taking conventions
- Working memory and processing speed (implicitly, through timed formats)
Note what is missing from that list: creativity, persistence, collaborative problem-solving, communication of complex ideas over time, and the ability to apply knowledge to novel real-world problems. These are not failures of standardized tests — they are features of what standardized tests were designed to do.
Demographic Gaps in Standardized Test Scores
The persistent demographic gaps in standardized test scores are among the most extensively documented findings in education research. On the SAT, the 2023 College Board data showed mean score gaps of approximately 177 points between white and Black students and 133 points between white and Hispanic students. These gaps have persisted for decades with relatively modest changes.
The causes are debated. Researchers distinguish between tests that are biased (where items systematically favor one group regardless of underlying skill), tests that reflect real differences in exposure to tested content (opportunity-to-learn gaps), and tests that accurately reflect skill differences that themselves result from systemic inequities. Most current evidence suggests the dominant mechanism is opportunity-to-learn: students who attend better-resourced schools, receive test preparation, and are exposed to academic language and content from an early age perform better — and these experiences are unequally distributed by race and income.
This matters for how parents should interpret scores. A standardized test score tells you a lot about a student’s exposure to tested content and academic language. It tells you less about their underlying intelligence or potential.
What Portfolio Assessment Actually Measures
What Is Portfolio Assessment?
Portfolio assessment, also called authentic assessment, asks students to compile evidence of their learning over time — completed projects, drafts, reflections, creative work — often with student commentary explaining what they learned and how their thinking changed. The assessment is typically based on a rubric that evaluates the work itself rather than a single-sitting performance under time pressure.
Project-based learning environments (see our deeper look at project-based learning research) are natural homes for portfolio assessment, but portfolios can exist in traditional classroom settings as well.
What does portfolio assessment capture? When implemented well:
- Growth over time (a portfolio shows where a student started and where they arrived)
- Creative and applied problem-solving (projects require applying knowledge to open-ended problems)
- Sustained effort and revision (portfolios often include multiple drafts)
- Student voice and metacognition (self-reflection is a core portfolio component)
- Communication of complex ideas in multiple forms (writing, visuals, presentations)
The 1990s saw a significant push for portfolio assessment in several states, including Vermont and Kentucky. The Vermont Writing Portfolio Program, studied by Koretz et al. (1994) in the Journal of Educational Statistics, found that portfolios measured dimensions of writing skill that standardized tests missed — particularly voice, organization of complex arguments, and revision quality. However, the same study found serious reliability problems: rater agreement was substantially lower for portfolios than for standardized assessments, meaning the same portfolio could receive very different scores from different evaluators.
The Reliability Problem
This is the central empirical challenge for portfolio assessment. Standardized tests, almost by definition, produce reliable scores — the same student taking the same test twice (absent coaching) will score similarly. Portfolio scores depend on rater judgment, and rater reliability in large-scale portfolio assessment has been consistently problematic.
This is not a fatal objection. Medical diagnosis, legal judgments, and performance evaluations all rely on expert human judgment that is less perfectly reliable than a scan or a number. The question is whether the reliability is sufficient for the purpose. For high-stakes decisions (graduation, college admission, gifted identification — see our article on how gifted identification works and its problems), the reliability bar is higher than for classroom feedback.
Equity Implications: A Different Set of Problems
Portfolio assessment is often advocated as a more equitable alternative to standardized tests. The equity picture is more complicated than this framing suggests.
Standardized tests: opportunity-to-learn gaps drive score gaps. Students from lower-income families score lower on average not because the tests are biased in the technical measurement sense, but because the content and language the tests assess are less familiar. Better-resourced schools provide more exposure to this content. The gap is real and reflects real differences in educational experience — but attributing those differences to student potential is a category error.
Portfolio assessment: implementation quality gaps drive outcome gaps. Well-implemented portfolio assessment requires thoughtful assignment design, clear rubrics, time for student revision, and teacher training in portfolio evaluation. These are not evenly distributed. Schools with experienced, well-supported teachers and smaller class sizes implement portfolio assessment much more effectively than under-resourced schools with high turnover and larger classes. A student at an under-resourced school with a poorly designed portfolio program may receive neither the feedback nor the accurate assessment that the approach promises.
This creates a troubling equity dynamic: advocates of portfolio assessment as an equity tool often underestimate that its equity depends entirely on implementation quality — which is itself unevenly distributed.
What Hybrid Approaches Look Like in the Data
Some school systems have moved toward hybrid models that use standardized tests for accountability and comparability while using portfolio or authentic assessment for learning and growth within classrooms. The research on these approaches is more encouraging than research on either approach in isolation.
New Hampshire’s Performance Assessment of Competency Education (PACE), which began in 2015, uses locally developed performance tasks (similar to portfolio assessment) for most student evaluation but includes standardized assessments as a comparability anchor. Early research from the National Center for the Improvement of Educational Assessment found that PACE schools maintained academic performance on external measures while reporting improved student engagement and teacher assessment literacy.
At the college level, the University of California’s admissions experiments have provided useful data. UC’s own research found that SAT/ACT scores added predictive validity beyond high school GPA — particularly for first-generation and underrepresented minority students whose high school GPA might reflect grade inflation or context that admissions officers cannot easily interpret.
Portfolio vs Standardized Assessment: What Each Measures
| Dimension | Standardized Tests | Portfolio/Authentic Assessment | Equity Implications | Best Use Case |
|---|---|---|---|---|
| Declarative knowledge | Measures well | Indirect evidence only | Gaps reflect opportunity-to-learn disparities | Comparability across schools; accountability |
| Mathematical computation | Measures well | Rarely assessed directly | Score gaps by income and race | Diagnosing skill gaps; curriculum alignment |
| Creative problem-solving | Does not measure | Measures well (when implemented well) | Depends on implementation quality | Classroom learning; project-based programs |
| Growth over time | Does not measure | Core strength | Implementation gaps disadvantage under-resourced schools | Formative feedback; documenting progress |
| Metacognition and reflection | Does not measure | Measures when rubrics include self-assessment | Teacher capacity affects quality | Student ownership of learning |
| Cross-disciplinary application | Limited | Measures when projects are integrative | Requires project design expertise | STEM and humanities integration |
| Reliability across raters | Very high | Low to moderate (requires rater training) | Reliability gaps disadvantage students without experienced teachers | High-stakes decisions require standardized anchors |
| Predictive validity for college GPA | Moderate (r ≈ 0.35–0.50) | Insufficient large-scale data | UC research shows GPA + test scores predict better than either alone | College admissions (as one factor among many) |
What to Watch for Over the Next 3 Months
- States expanding test-optional graduation requirements: Several states are revisiting standardized test graduation requirements post-pandemic. Watch for implementation details — the equity implications depend heavily on what replaces tests.
- College admissions test-optional policy reviews: Many selective universities promised to revisit test-optional policies after 3–5 years. 2026 is the year several of those reviews are due. MIT and Yale have already returned to requiring scores; others are still deciding.
- New Hampshire PACE expansion: The most rigorous hybrid assessment experiment in the U.S. is releasing 5-year outcome data. This will be the best evidence yet on whether authentic assessment can be implemented at scale without sacrificing accountability.
- AP and IB portfolio components: Both Advanced Placement and International Baccalaureate programs have expanded portfolio and performance task components in recent redesigns. Watch for outcome data comparing students who did and did not use these features.
Frequently Asked Questions
Are standardized tests biased against students of color? This is a question that requires careful distinction. In the technical measurement sense, most standardized tests in common use have been examined for item-level bias and items that show differential functioning are typically removed. But the tests reflect educational experiences that are unequally distributed by race and income — so the scores reflect real opportunity gaps. Whether you call that “bias” depends on your definition. The score gaps are real; the question is what causes them and what to do about it.
Can portfolio assessment replace standardized tests for college admissions? Not easily, for the reason that standardized tests were invented: they provide comparability across very different contexts. A portfolio from a student at an elite private school and a portfolio from a student at an under-resourced public school are extremely difficult to evaluate on the same scale. Standardized scores, for all their problems, allow colleges to make cross-context comparisons. This is why most selective colleges, even test-optional ones, still consider standardized scores when students submit them.
What does research say about whether kids learn more in portfolio-based classrooms? The research is mixed and depends heavily on implementation. Well-implemented project-based learning, which often incorporates portfolio assessment, shows positive effects on student engagement and transfer of knowledge to novel problems. Poorly implemented versions show no advantage over traditional instruction. The effect is real but contingent on teacher skill and program design.
How do standardized test scores interact with grades in predicting success? Multiple large studies find that high school GPA and standardized test scores together predict college outcomes better than either alone. This supports a complementary view: grades reflect effort, work habits, and accumulated knowledge across many assignments; tests reflect a specific type of academic skill performance. Both are real, and they capture partially different things.
Are there schools where portfolio assessment works well at scale? Yes. Schools in the Coalition of Essential Schools network, schools using the Expeditionary Learning/EL Education model, and International Baccalaureate schools have implemented portfolio and performance assessment at scale with documented positive outcomes. These schools tend to have strong professional development infrastructure, low teacher turnover, and institutional commitment to assessment quality.
What should parents ask when evaluating a school’s assessment approach? Key questions: What is the school’s approach to tracking individual student growth over time? How does the school communicate about areas where a student needs additional support? How do students’ outcomes on external measures (state tests, AP exams) compare to schools with similar demographics? What is the teacher turnover rate — a proxy for the instructional consistency that portfolio assessment requires?
How do standardized test scores predict long-term outcomes beyond college GPA? The predictive validity of standardized tests weakens significantly beyond first-year college GPA. Long-term earnings, job performance, and life satisfaction show much weaker correlations with standardized test scores than with non-cognitive skills like persistence, collaboration, and communication — which portfolio assessment is better positioned to develop and measure.
Are there equity-minded ways to use standardized tests? Yes. Using standardized tests to identify students who are performing well on external measures but underperforming in classroom grades (a potential sign of unrecognized ability) is one equity-positive use case. Using test data to identify schools and districts where achievement gaps are largest and resources are most needed is another. The equity problem is not with tests per se but with treating test scores as fixed measures of potential rather than as measures of current exposure to tested content.
About the Author
About the author Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.
Sources
- Koretz, D., Stecher, B., Klein, S., & McCaffrey, D. (1994). “The Vermont Portfolio Assessment Program: Findings and Implications.” Educational Measurement: Issues and Practice, 13(3), 5–16.
- College Board. (2023). “SAT Suite of Assessments Annual Report.” collegeboard.org.
- Geiser, S., & Santelices, M. V. (2007). “Validity of High-School Grades in Predicting Student Success Beyond the Freshman Year.” Center for Studies in Higher Education, UC Berkeley.
- National Center for the Improvement of Educational Assessment. (2021). “New Hampshire’s PACE: An Update on the First Five Years.” nciea.org.
- Au, W. (2007). “High-Stakes Testing and Curricular Control: A Qualitative Metasynthesis.” Educational Researcher, 36(5), 258–267.
- Darling-Hammond, L., & Adamson, F. (2010). Beyond Basic Skills: The Role of Performance Assessment in Achieving 21st Century Standards. Stanford Center for Opportunity Policy in Education.
- University of California. (2020). “Standardized Testing and the University of California: A Review.” universityofcalifornia.edu.