Showing posts with label New York City report cards. Show all posts
Showing posts with label New York City report cards. Show all posts

Wednesday, December 19, 2007

4 Good Ones on NYC Report Cards

Four useful articles on report cards:

1) Diane Ravitch's NY Sun op-ed from earlier this week.
2) Leonie Haimson's post at NYC Parents, which provides links to youtube footage from the City Council hearings.
3) Norm Fruchter's Ed Week commentary.
4) Sam Freedman's NYT column.

My prior posts on report cards are archived here. Basically, here's what they say: even if your only concern is statistical validity, the NYC report cards are a mess. The system doesn't take into account measurement error, creates unreasonable peer groups, and pays no attention to regression to the mean (the idea that very high or low data points will move toward the mean the next time they are measured - i.e. Staten Island's PS 35, discussed in Randi Weingarten's commentary). And that's just a small sampling of the design problems.

What the report card debate reveals, I think, is that parents and citizens don't view school quality as a unidimensional construct. Nor do all parents want the same things from their schools; or, more accurately, they may want the same things, but put different weights on academic growth, climate, etc. Some value overall performance more than growth. Others care more about the climate and social environment than academics. Still others prioritize safety. In short, assigning one grade conflicts with our competing intuitions about how to value different dimensions of schooling.

If a central goal of the report card system is to provide information to parents, the Dept of Ed should consider assigning multiple grades. First, though, the basic design of the system needs some serious work.

Monday, December 10, 2007

If only I had my invisibility cloak...

Philissa at insideschools reports that Jim Liebman could have used an invisibility cloak at the City Council Report Card hearings today. Patrick Sullivan posts additional details, and the NYT weighs in here.

Sounds like a youtube sensation in the making. Cloaks and jokes aside, see my prior posts about NYC report cards for more serious insights (archived here).

Wednesday, December 5, 2007

Imperial Death March Sounds for 6 NYC Schools

The NYC Department of Education announced that it will phase out 6 schools this year serving 3417 students. NY1 reports that up to 14 additional schools may be closed by the end of this year.

What's perplexing, though, is why these six schools were chosen - 3 of the 6 got Ds, not Fs. If you really believe your progress report system identifies "failing" schools, why would you close D schools before F schools? Note to Joel: legitimate regimes follow their own rules.

Some basic facts about the schools that are slated to close:
  • 3 are in Manhattan, 2 are in the Bronx, and 1 is in Brooklyn.

  • 3 are middle schools, 1 is a K-8, 1 is an elementary school, and 1 is a high school.

  • All schools serve a high proportion of free lunch eligible students (ranging from 66.3 to 86.4%).

  • Five of the six schools serve 60% or more Hispanic students.

  • Two of the six schools are more than 25% ELL.

  • Four of the six schools have very high concentrations of special education students (ranging from 17.1 to 26.1%).
See the two tables below for more details.

Tuesday, November 20, 2007

Overheard in New York


After almost five and a half years as chancellor, I know you can’t point to a single number, be it a test score or graduation rate, to prove success or failure. The whole picture is important.

--Joel Klein

Unless, of course, you're talking about grading schools - where you can just point to a single letter.

In a memo, Klein goes on to say:

We need to look at all the indicators we have—promotion rates, graduation rates, State English and math test results, Regents pass rates, national test results, and even the results of things like Advanced Placement and PSAT exams—to understand how well our students are performing and progressing. Together, these factors paint a complete picture. In that context, I am encouraged by the results we got last week on a key national indicator, the National Assessment of Educational Progress results.

Monday, November 19, 2007

NYC Report Card Haiku Finale!

Last week, readers wrote 68 haiku, which are published here as the New York City Report Card Haiku Magazine! Read it, print it, share it.

In lieu of holding a vote for the best haiku, I've decided to share my five favorite ones instead:

Amateur Night's spozed
to be at the Apollo
not at Tweed courthouse
-Anonymous 7:50 AM

Step onto the scale
I'm fatter than my neighbor
So I'll get a C
-Anonymous 12:31 AM

My school got an A
I never learned to Haiku
I read passages
-Doug Douglass

The wise little tree
Bends with the strong winds and does
Not break easily
-Anonymous 10:21 PM

The only numbers
that level the playing field
have a dollar sign
-Anonymous 10:54 PM

Thank you to everyone who contributed a haiku!

Monday Morning NYC Report Card Factoid

NCLB uses a very different metric of school success than the NYC Report Cards. What percentage of A-F schools in NYC are schools in need of improvement under NCLB?

  • Of A schools, 17.9% (50 schools) are schools in need of improvement under NCLB.

    Of B schools, 29.8% (139 schools) are schools in need of improvement under NCLB.

    Of C schools, 32% (101 schools) are schools in need of improvement under NCLB.

    Of D schools, 32.3% (32 schools) are schools in need of improvement under NCLB.

    Of F schools, 32% (16 schools) are schools in need of improvement under NCLB.

    Of schools with grades still "under review" as of today, 50% (7 schools) are schools in need of improvement under NCLB.

What's a principal to do? To succeed under NCLB, schools need to lift kids over the proficiency bar. To succeed in NYC, schools need to focus on the growth of the lowest performing kids.

On Thursday, you can be thankful for turkey, your family, and the fact that you are not a principal in NYC.

Friday, November 16, 2007

The Trouble with NYC Report Card Peer Groups

Central to the NYC Report Cards is the idea of a "peer group." The concept is a good one - it's not fair to compare schools serving vastly different populations of kids, as does NCLB, because kids may be doing worse in school because of out-of-school issues the school can't control.

The NYC DOE deserves credit for moving away from these apples-to-oranges comparisons. But NYC constructed elementary school and K-8 peer groups using a weighted formula of only four school characteristics - percent black and Hispanic combined, percent free lunch, percent special ed, and percent ELL. The junior high and high school peer groups were constructed using 4th and 8th grade scores, respectively. Each school is then assigned a "peer index" number, and the 20 schools above and below a school in the peer index serve as their "peer group."
How did this work out? Consider elementary school PS 196 in Queens, which received a B, as an example. PS 196 serves students who are:
  • 40% Asian, but is being compared to schools that have between 1% and 69% Asian students (average=17.8%).
  • 44.8% white, but is being compared to schools that have between 22.1% and 90% white students (average=66.6%)
  • 2.1% African-American students, but is being compared to schools that have between 0% and 16.7% African-American students (average=3.96%).
  • 13.1% Hispanic students, but is being compared to schools that have between 2.9% and 17.1% Hispanic students (average=10.4%)
  • 14.6% free lunch students, but is being compared to schools that have between 2.4% and 26.8% free lunch students (average=15.5%).

Key inputs that are not in control of the school vary widely as well:

  • PS 196 is at 106.4% of building capacity, but is being compared to schools that are between 51.6% and 136.5% of capacity (average=92.1%)
  • 54.3% teachers at PS 196 have more than 5 years experience, but is being compared to schools that have between 35.7 and 81.3% teachers with more than 5 years experience (average=59.5)
  • PS 196 has 658 students, but is being compared to schools that have between 180 and 1263 students (average=601).

The comparison groups were constructed differently at the high school level, where middle school test scores, not demographics, were used to create comparison groups. But do middle school test scores provide enough information to net out all factors that influence students' graduation?

Consider the Manhattan Center for Science and Mathematics, a selective school located in East Harlem that screens students not only on their test scores, but on their middle school grades and attendance. However, the Manhattan Center is not only compared with other schools that select students on their test scores, grades, and attendance; in fact, a handful of high schools with zoned programs are included in its peer group.

These schools have students with widely varing characteristics as incoming 9th graders. The Manhattan Center serves students who have the following characteristics as entering 9th graders:

  • 61.2% proficient in reading in 9th grade, but is being compared to schools that have between 37.5 % and 100% proficient as entering 9th graders (average=70%) .
  • 79.5% proficient in math in 9th grade, but is being compared with schools that have between 45.2% and 100% proficient in 9th grade (average=77%)
  • 65% free lunch students, but is being compared to schools that have between 12.2% and 93.9% free lunch students (average=39.8%)
  • 16.8% of students who are overage for grade (i.e. they have been held back), but is being compared to schools that have between 3.3% and 61% overage for grade (average=12.5%).
  • 12.4% ELL, but is being compared to schools that have between 0 and 84.1% ELL (average=5.01%)
  • 7.1% full-time special education, but is being compared to schools that have between 0 and 4.7% full-time special education (average=.8%)

There are two problems with this approach - one statistical and one practical:

1) As the comparisons above demonstrate, the peer groups falsely provide the illusion of fair comparison.

2) Perhaps more important for responses to these report cards - any educator who sits down with numbers will likely conclude that these comparisons lack face validity. And if these peer groups are not believeable to educators - that is, they don't feel that they have an equal chance of winning or losing this game - we are going to see enormous amounts of playing the system as schools attempt to succeed in a system that is perceived as fundamentally unfair.

Last Chance for NYC School Report Card Haiku Contest

our writers' workshop
activism by haiku
start a file on her

Yesterday, I set the goal of at least 6o haiku by the end of the day on Friday, and now we have 58! Push us over the edge. As promised, I'll format them as a lit mag so we all have a souvenir.

The haiku contest runs through the end of the day today (Friday) - I'll pick three I like and we'll vote next week. The only rules are the 5-7-5 syllable one, and to keep it playful and G rated.

Here are some submissions from yesterday:

The only numbers
that level the playing field
have a dollar sign
-Anonymous 10:54PM

Who would have thunk it?
A system that is worse than
No Child Left Behind
-Anonymous 10:38PM

It's only Kool-Aid
A little sip won't hurt you
said the Chancellor
-Anonymous 11:16PM

Monday, November 12, 2007

NYC Parent and Teacher Surveys: Who Responded?

Last spring, the NYC Dept of Education spent $2 million administering a survey to all parents, teachers, and students in the school system. What did the response rates look like by school, and did these response rates vary in a patterned way? Here’s what I found:

1) Parents and teachers at A schools were more likely to respond to surveys than parents and teachers at F schools. “A” schools had an average parent response rate of 31.75% versus 22.59% at F schools. The differences are relatively small for teachers: 48.23% versus 45.03%.



There’s something in this finding for everyone. One could argue that A schools are actually higher quality schools and more parents are involved as a result, which is why a higher proportion of parents responded. Alternatively, it may be that schools with lower levels of parent support are struggling for this very reason. The graph below, which plots the overall score for high schools against the parent response rate, provides some support for this point (correlation=.54).


2) Schools with higher proportions of Hispanic, African-American, and free lunch kids had much lower parental survey response rates, while those with higher proportions of White and Asian kids had much higher response rates.



To make the graph above, I divided schools into quintiles – five equally sized groups that each represent 20% of the schools. Schools in Quintile 1 have the lowest proportions of a given group, while schools in Quintile 5 have the highest proportions of that group. For example, Quintile 1 schools for the free lunch population have 48.7% free lunch or fewer, while Quintile 5 schools have 85.25% free lunch kids or more. For the Hispanic population, schools with 14.2% Hispanic kids or fewer are Quintile 1 schools, while those with 65.35% Hispanic kids or more are Quintile 5 schools.

3) This relationship holds for the teacher survey response rates. Schools with higher proportions of African-American and free lunch kids had significantly lower teacher survey response rates, while those with higher proportions of White and Asian kids had higher response rates.

Overall, these results raise the question of the validity of the parent surveys for any given school. 25% of schools had parent response rates of 18% or less (one school had a parent response rate of 1.97%!).

In addition, given how much these results vary by demographics, it is also not clear that we can validly compare actual survey responses across all schools. For example, is it meaningful to look at the safety and respect scores of a school with a parent response rate of 67% and compare those with a school with a response rate of 9%?

Lots of food for thought here – as always, email me if you'd like to see the full tables.

Sunday, November 11, 2007

NYC School Report Card Groundhog Week


Last week, I wrote about NYC School Report Cards. (All posts archived here.) Though I was planning on looking at abstinence only sex education this week, there's a lot left to say about report cards. So here we are again. This week:

Monday: New York City Parent and Teacher Surveys: Who Responded?

Friday: Peer Comparison Schools - Now that the peer indices have been released, we can take a closer look.

In Progress: Extra Credit - Who Got It?

Friday, November 9, 2007

Is this a wake-up call for the people who work there? You betcha.

Earlier this week, Mayor Michael Bloomberg flexed his muscles by threatening to close F schools as early as June. He quipped, "Is this a wake-up call for the people who work there? You betcha."

Through analyzing these data, I've concluded that the people in need of a wake-up call work not at F schools, but at the NYC Department of Education. Undoubtedly, data can and should be used for organizational learning and school improvement. But if we're going to rank and sort schools - an action that has serious consequences for the kids, educators, and parents affected - the Department of Ed's methods should be in line with the standards to which statisticians and quantitative social scientists hold themselves. Needless to say, NYC's report cards are not.

There are five reasons the report cards might kindly be called statistical malpractice:

1) Ignoring measurement error

Measurement error isn't sexy and won't attract the attention of journalists and commentators. But it may be the central downfall of the NYC report card system. For example, elementary school PS 179 (score=30.9) got an F, while PS 277 (score=31.06) got a D. Similarly, Queens Gateway to Health Sciences Secondary School (score=65.21) got a B, while IS 229 (score=65.22) got an A.

If we actually acknowledged that these overall scores are measured with error, a school scoring a 65.21 is not statistically distinguishable from one scoring a 65.22 (a difference of .0007 standard deviations) . And Mayor Bloomberg is threatening to close PS 179 this year and keep PS 277 open because of a difference of .16? (See the grade brackets in the table below to see how close your school was to earning a higher or lower grade.)


2) Arbitrary grade distributions and cutoffs

Initially, the Dept announced a curve on which schools would be graded, but now they've curiously changed the distribution of grades and created different distributions for elementary, middle schools, high schools, and K-8s. Why should 25.26% of middle schools get As, but only 21.72% of elementary schools do? By the same token, why should 5.12% of high schools get Ds while 9.69% of middle schools do? It's not that the Dept has set criterion-referenced score cutoffs for attaining these grades, as the table above demonstrates - so what's going on here? The Dept of Ed needs to release more information about why this distribution of grades was chosen, and why it is different for each school level.

For example, there are more A/B middle schools than A/B elementary schools - does this mean that NYC's middle schools are "better?" The table below shows the percentage of schools receiving each grade for each school level, as well as the number of schools receiving that grade. You can click to enlarge.

* For high schools, the denominator does not include schools with grades "under review."

3) 6-12 schools grade discrepancies

Schools serving grades 6-12 got two grades - one for 6-8, and one for 9-12. These are the same schools, same principal, same roof. But for the 33 schools for which there are middle and high school grades available, 22 have different grades - though they are the same school! Sometimes these differences are substantial.

Consider the Academy of Environmental Science - its high school got a C, but its middle school got an F. At Hostos Lincoln Academy of Science, the middle school got a D, but the high school got a B. At the Bronx School for Law, Government, and Justice, the middle school got an F, but the high school got a C.

4) Poorly constructed comparison groups

As I've written here, the Dept of Ed flubbed the comparison groups by treating the percent African-American and percent Hispanic as interchangeable (i.e. a school with 59% Hispanic and 1% African-American is a perfect match for a school with 59% African-American and 1% Hispanic.) In addition, the Dept did not consider the proportion of Asian students when creating comparison groups; schools with higher proportions of Asian kids were more likely to get As and Bs, and there's no reason to believe that Asian kids in NYC have access to much higher quality schools. It's more likely that Asian kids grow academically at a faster rate because of things that happen outside of school.

Until the Dept releases the comparison groups, it is difficult to know how bad these comparisons are - so stay tuned.

5) Problems with growth models: Interval scaling and ceiling effects

I'm all for growth models, but you can't treat 1 unit of growth at the bottom of the distribution (i.e. moving from 13 to 14 on a 100 point scale) the same as 1 unit of growth at the top of the distribution (i.e. moving from an 89 to a 90). Put formally, the Department of Ed's model assumes that tests are "interval scaled," but they are not. Similarly, if a student is scoring near the top possible score of a test (the ceiling), there is very little room left to grow. One can address this problem by weighting growth at different parts of the distribution differently, but the Dept chose not to do this.

Hopefully, folks who care about public education in NYC will issue a wake-up call to the Dept of Ed and demand that these problems are fixed before vital decisions are made about schools based on highly questionable methods.

Thursday, November 8, 2007

What Does It Mean for a School to Be Good? More on NYC Report Cards

I promised a post on the theory behind assigning schools a single letter grade, so here we are.

The idealized idea is that in the vast supermarket of educational goods, grades serve as a strong signaling mechanism that tell both educators and parents where their schools stand. They are intended to provide a stronger signal, and thus clearer information to guide action, than regular test scores. As a result, their proponents argue that by placing schools in a distribution where everyone knows the pecking order, grades create strong incentives for improvement. They are also supposed to serve as shortcuts for parents, which allow them to vote with their feet and create market pressure for schools to improve.

In the locales where grades have been implemented, they appear to have real consequences on behavior. For example, economists David Figlio and Maurice Lucas found that trivial differences in grades in Florida affect housing prices. Similarly, a new experimental study by economist Justine Hastings and colleagues found that providing simplified information on test scores "improves" parents' choices ( i.e. they pick higher scoring schools).

All of this makes sense if school quality is unidimensional - but as Diane Ravitch has pointed out, there are many dimensions of schooling. Does it make sense to give an overall grade, when schools may excel at some dimensions and not others? Parents look for different things in schools, and while some may prioritize a positive school environment, others may care more about value-added. Still others may care about overall proficiency rates. I took a look at the school environment and Quality Review scores in the report cards, which provide insight about the general environment of the school. This is not to say these measures are great, but they're something to look at.

What I found is that an A or B school is not necessarily a school with a positive school environment or a well-developed rating on the Quality Review, and vice versa:

* 67 schools that received school environment scores in the lowest 20% of all city schools (i.e. if this was a separate grade, they would have received an F) received As and Bs.

* On the other hand, 8 schools that were in the highest 20% of environmental category scores received Ds or Fs.

* 35 schools that were rated as "Underdeveloped" on their Quality Reviews received As or Bs, while 22 schools rated as "Well-developed" received Ds or Fs.


New Yorkers may be wondering about which schools fell into these mismatch categories:

Schools that got As but were in the bottom 20% of school environment category scores:

1) PS 165 ROBERT E SIMON
2) JHS 080 THE MOSHOLU PARKWAY
3) BRONX LITTLE SCHOOL
4) PS 038 THE PACIFIC
5) PS 202 ERNEST S JENKYNS
6) PS 015 JACKIE ROBINSON
7) PS 031 BAYSIDE

Schools that got Ds or Fs but were in the top 20% of the school environment category scores:

1) MS 243 CENTER SCHOOL
2) CENTRAL PARK EAST I
3) ISAAC NEWTON JHS FOR SCIENCE & MATH
4) PS 291
5) PS 007 ABRAHAM LINCOLN
6) BROOKLYN COLLEGE ACADEMY
7) PS 179
8) PS 35 The Clove Valley School

Schools rated as "underdeveloped" (the lowest category) on Quality Reviews that got As and Bs:

As:

1) MS 260 CLINTON SCHL WRITERS &
2) THE EAST VILLAGE COMMUNITY SCHOOL
3) NEW EXPLORATIONS INTO SCIENCE, TECHNOLOGY AND MATH
4) THE HERITAGE SCHOOL
5) JHS 080 THE MOSHOLU PARKWAY
6) SCHOOL FOR EXCELLENCE
7) JHS 050 JOHN D WELLS
8) PS 193 GIL HODGES

Bs:
9) PS 015 ROBERTO CLEMENTE
10) JHS 044 WILLIAM J O'SHEA
11) PS 137 JOHN L BERNSTEIN
12) PS 217/IS 217 ROOSEVELT IS.
13) NEW EXPLORATIONS INTO SCIENCE, TECHNOLOGY & MATH
14) PS/IS 54
15) PS 061 FRANCISCO OLLER
16) PS 072 DR WILLIAM DORNEY
17) PS 107
18) I.S.117 JOSEPH H WADE
19) JHS 125 HENRY HUDSON
20) JHS 162 LOLA RODRIGUEZ DE TIO
21) PS 246 POE CENTER
22) NEW MILLENNIUM BUSINESS ACADEMY MIDDLE SCHOOL
23) ACCION ACADEMY
24) MS 391
25) HIGH SCHOOL FOR TEACHING AND THE PROFESSIONS
26) GLOBAL ENTERPRISE HIGH SCHOOL
27) PS 046 EDWARD C BLUM
28) PS 133 WILLIAM A BUTLER
29) IS 136 CHARLES O DEWEY
30) PS 184 NEWPORT
31) ANDRIES HUDDE
32) PS 305 DR PETER RAY
33) PS 040 SAMUEL HUNTINGTON
34) I.S. 73 - THE FRANK SANSIVIERI INTERMEDIATE SCHOOL
35) PS 156 LAURELTON

Schools graded as well-developed (the highest category) on Quality Reviews that got Ds and Fs:

Ds:

1) MS 243 CENTER SCHOOL
2) PS 016 WAKEFIELD
3) PS 018 JOHN PETER ZENGER
4) PS 044 DAVID C FARRAGUT
5) PS 047 JOHN RANDOLPH
6) PS 007 ABRAHAM LINCOLN
7) I. S. 381
8) BROOKLYN COLLEGE ACADEMY
9) JHS 008 RICHARD S GROSSLEY
10) PS 105 THE BAY SCHOOL
11) PS 166 HENRY GRADSTEIN
12) PS 8 SHIRLEE SOLOMON
13) I S 034 TOTTENVILLE
14) PS 041 NEW DORP
15) PS 055 HENRY M BOEHM

Fs:
16) PS 033 CHELSEA PREP
17) PS 106 PARKCHESTER
18) PS 130 ABRAM STEVENS HEWITT
19) PS 179
20) PS 182
21) PS 238 Anne Sullivan
22) PS 35 The Clove Valley School

So what does it mean for a school to be good? As saavy parents and teachers know, it depends on what "good" means. The report card grades should be interpreted with that in mind.

Wednesday, November 7, 2007

NYC School Report Cards: How Did the Community School Districts Fare?

Here are a few figures on how the Community School Districts fared in the report cards. More commentary and interpretation on this later. You can click on the table below to enlarge it.




NYC School Report Cards II: A Closer Look

One of the most striking patterns from my post yesterday was that schools with higher report card grades, on average, had higher proportions of Asian kids than those with lower grades. This raised serious questions for me about the validity of the comparison groups, which were generated using only four school characteristics: the combined African-American and Hispanic population, percent free/reduced lunch, percent special ed, and percent ELL.

From national datasets like the Early Childhood Longitudinal Study, it's clear that Asian kids have different growth trajectories throughout elementary school. It's also clear that Hispanic and African-American kids have very different growth trajectories, with African-American kids falling behind at a much faster rate. If a large proportion of the report card is based on growth, but the growth measures don't account for varying trajectories that are, in part, the result of non-school factors, schools serving higher proportions of Asian kids will look like they're producing more growth. Similarly, schools serving higher proportions of Hispanic kids will look like they are producing more growth if they are compared to schools serving similar proportions of African-American kids. [We can debate how much of these differences in trajectories are explained by school versus outside-of-school factors.]

Here's what I found (I'll interpret a whole line of the table so I'm clear on what these measures are):

  • At the elementary school level, A schools have an average of 20.99% Asian students. By contrast, the median A school has 9.55% Asian students. This tells us that the average is pulled up by schools with very high proportions of Asian kids. The standard deviation is there for stats junkies, but most readers will prefer to look at the interquartile range (the 25% and 75% columns) to get a sense of how much variation there is. If we read across the A row, the 25% column tells us that 25% of A schools have 2.3% Asian or fewer. Similarly, the 75% column says that 25% of A schools have 36.8% Asian students or more. The range column represents the lowest and highest values for a given grade; A schools have between .2 and 92.6% Asian students.

  • If we compare medians instead of averages, we still see that A elementary schools have more than 3 times as many Asian students than F schools (9.55% versus 2.90%). The 25% of A schools with the highest Asian population have between 36.8% and 92.6% Asian students, while the 25% of F schools with the highest Asian populations have 5.3% Asian students and 28.5% Asian students.

  • The general pattern is similar across all levels of schooling; A schools have substantially higher proportions of Asian students.

The tables below show the distribution of Asian students by school grade and school level (i.e. elementary, middle, K-8, and HS):

* 3 elementary schools had missing data in 2005 and are thus not represented in the table above.


* Because many 6-12 schools opened recently and thus do not have 2005 data available, the middle school data are missing 32 schools, including 10 A schools, 14 B schools, 5 C schools, 1 D school , and 2 F schools, and thus should be interpreted with this in mind.


* 1 K-8 school had missing data in 2005 and is not represented in the table above.
* 1 high school had missing data in 2005 and is not represented in the table above.

Unless we believe that schools with high concentrations of Asian students are higher quality schools, these results raise a lot of questions about the validity of these comparison groups. In the multivariate analyses that I describe below, I also find that schools with higher proportions of Hispanic students are more likely to receive A or B grades - which raises the question of the accuracy of using an aggregate black/Hispanic number to compare schools.

Overall, these results suggest that the Dept of Ed's method of establishing comparisons groups ultimately results in apple-to-orange comparisons.

For geekier analyses, read on:

For those who are interested, I also ran a series of descriptive logistic regressions for the purpose of examining the association between school racial composition and schools' odds of earning an A or B grade, net of many other factors that could explain this association. In these models, I controlled for percent free lunch, percent female, percent immigrant, percent stability (i.e. the opposite of mobility), percent full and part-time special education, percent ELL, school size, percent capacity (how crowded the school is), and teacher characteristics (percent with more than 5 years teaching and percent with a masters degree). I didn't impute missing values, so these analyses include 970 schools of the 1187 that had data available in 2005. Remember, regressions like these are just descriptive, not causal [that is, they describe patterns observed in the data and do not necessarily explain *why* a school received the grade it did]. Nonetheless, unless we believe that schools with the highest concentrations of Hispanic and Asian students are much higher quality than those with lower concentrations of these students, these results suggest that the peer comparison groups are not entirely fair:

  • First, I divided schools into four equal groups - quartiles - based on their percent Asian. (This is a typical approach to modeling non-linearities.) The first quartile included the 25% of schools that have the lowest proportions of Asian students, and the fourth quartile included the 25% of schools that have the highest proportions of Asian students - in these analyses, quartile 4 includes >15.3% Asian. I did the same thing with the African-American and Hispanic populations, i.e. divided schools into four quartiles.
  • When we just examine the association between racial composition and a schools' odds of getting an A or a B, quartile 4 schools (those with the most Asian students) are almost 2.5 times more likely to get an A or B (odds ratio=2.44, p=.001) compared to those with the lowest proportions of Asian students (those with 1.5 percent Asian or less) . However, schools in the 2nd and 3rd quartiles of the Asian population have no advantage over those in the first. When we just look at the relationship between getting and A/B and African-American and Hispanic composition, we see no stastically significant relationship.
  • Once we control for all variables listed above, quartile 4 Asian schools have a smaller advantage (they are slightly less than twice as likely to get an A or B (odds ratio=1.87, p=.041). Again, there is no advantage of being in the 2nd or 3rd quartile of the Asian population over the first. But schools in quartile 4 of the Hispanic population have an even larger advantage - they are slightly more than two times as likely to get an A or B (odds ratio=2.12, p=.036).
  • Nonetheless, this full set of predictors only explains a tiny proportion of the variance (pseudo R2=.06).

If anyone is interested in checking out the full results, email me and I'll send you the output.

Tuesday, November 6, 2007

The NYC School Report Card: A First Cut at the Data

The report cards are out - you can check them out here. There's a lot to talk about, but I thought it would be helpful to first create a profile of the schools that received various grades. Using the progress report data and data from the 2005 release of the NYC School Report Cards, I looked at the average characteristics of schools receiving a given grade (i.e. means). For example, A schools average 15.92% Asian students, while schools that received Fs have an average Asian population of 4.53%.


The bright line findings from the table below are:
  • Schools with higher grades have higher proportions of Asian students; A schools have an average of 16% Asian, while F schools have 4.5%.
  • Schools with higher grades have much lower proportions of African-American students; A schools have an average of 26%, while F schools have an average of 47%.
  • Schools with higher grades have lower proportions of students qualifying for free and reduced lunch - (A schools=65%, F schools=76%). Interestingly, schools "under review" - those contesting their grades - have much lower percentages of free and reduced lunch kids (56.5%).
  • Schools with higher grades have higher proportions of Hispanic and immigrant students.
I'll write in more depth about the comparison group issue later, but for now, these results - particularly the Asian finding - suggest that the comparison groups may not adequately control for the fact that kids of different backgrounds may be on different growth trajectories for reasons that have little to do with the schools they attend - read more about this under point #2 here.


Also interesting are the dimensions where there are very few differences between schools receiving different grades, as you'll see in some of the variables below.

* Notes on the analysis: The most recent data to which I had access were 2005 data; because these variables are highly correlated across years, the general trends should look the same. (If someone has clean 2006 data, send it along and I can rerun these descriptives.) Also note that these are means; if you are interested in standard deviations and ranges, email me.

Monday, November 5, 2007

The Report Card Management Strategy

Five years ago, Malcolm Gladwell wrote an article called "The Talent Myth" questioning the management zeitgeist that the NYC Department of Education has swallowed wholesale:
At the heart of the McKinsey vision is a process that the War for Talent advocates refer to as "differentiation and affirmation." Employers, they argue, need to sit down once or twice a year and hold a "candid, probing, no-holds-barred debate about each individual," sorting employees into A, B, and C groups. The A's must be challenged and disproportionately rewarded. The B's need to be encouraged and affirmed. The C's need to shape up or be shipped out. [One company] followed this advice almost to the letter, setting up internal Performance Review Committees. The members got together twice a year, and graded each person in their section on ten separate criteria, using a scale of one to five. The process was called "rank and yank." Those graded at the top of their unit received bonuses two-thirds higher than those in the next thirty per cent; those who ranked at the bottom received no bonuses and no extra stock options--and in some cases were pushed out.
Gladwell writes at length about the management strategy of a dazzlingly successful company, which:
took more credit for success than was legitimate, that did not acknowledge responsibility for its failures, that shrewdly sold the rest of us on its genius, and that substituted self-nomination for disciplined management.
How did this work out?
The broader failing of McKinsey and its acolytes...is their assumption that an organization's intelligence is simply a function of the intelligence of its employees. They believe in stars, because they don't believe in systems. In a way, that's understandable, because our lives are so obviously enriched by individual brilliance. Groups don't write great novels, and a committee didn't come up with the theory of relativity. But companies work by different rules. They don't just create; they execute and compete and coördinate the efforts of many different people, and the organizations that are most successful at that task are the ones where the system is the star.
Gladwell concluded:
They were there looking for people who had the talent to think outside the box. It never occurred to them that, if everyone had to think outside the box, maybe it was the box that needed fixing.
What company was Gladwell writing about? Enron.

I hesitate to invoke the comparison here, if only because Enron has become a cognitive shortcut term for too many things. Here's the parallel - Enron's organizational failure was not just the result of a handful of swindlers, i.e. Lay and Skilling. Enron created a organizational structure and incentive system that virtually ensured that gaming and malfeasance would occur.

As we wait for the grades, check out the entire Gladwell article.

Sunday, November 4, 2007

Reporting on NYC's Report Cards

eduwonkette was not intended to be a blog about NYC school reform, but the folks up in NYC provide lots to write about. This week, the NYC Department of Education will release report cards grading its schools - see the NY Times article here. While 15% of schools will get As, 5% will get Fs. Here's this week's outline:

Monday - The Report Card Management Strategy in Other Sectors

Tuesday - NYC School Report Cards I: What kinds of schools received As, Bs, etc?

Wednesday: Did the Department of Education Flub the Peer Comparison Groups?: A Closer Look at the NYC Report Card Data

Thursday - What Does It Mean for a School to Be Good?: In this post, I explain the theory behind NYCs school grades, and look at schools that did well in their environmental category score but poorly overall, and vice versa.

Friday - Problems with the NYC Report Cards