rec.games.trading-cards.jyhad

Request for Comments (Doing Away With ELO Ratings)

42 messages from 20 participants · 24 January 2002 – 05 February 2002
original thread on Google Groups

The Lasombra

Yes, the current rating system is flawed. Terminally, I suspect. Take for example the #1 player in the world for 18+ months, one of the players from Vienna, Alex something or other. He played in 3 tournaments, won them all, then stopped playing in tournaments for almost 2 years. Was he really a better player than Rob Treasure all that time? I posit that he was not. Is Stefan Ferenci (with 3 tournaments) really a better player than Remy Auclair (with 14+ tournaments)? The ratings say he is. In one tournament, if you win your first round, ie Rob in Lafayette, then take 0vp in rounds two and three, you have a nice low rating going into the finals. When Rob then beat players that took time to warm up, say Jeff with 0,1,3 vp going into the finals, both player's ratings swings are inordinately huge. I would prefer to see some sort of earned runs average type of thing myself. For example, three or four columns, all additive, no real math for the maintainer: 1) Tournament Rounds played 2) Victory Points earned (alt. Game Wins earned) 3) Tourament Final appearances 4) Tournament Wins This both eases the burden on the maintainer, ie Todd, and gives a better picture of how a person is performing over time. This also allows tournament results to be entered and updated immediately, without concern that another tournament from 1 week earlier hasn't been received by the maintainer and having that hold up 30+ people's rankings. All the people in the top 100 that have only played in 3 tournaments (10-12 rounds) keep their high percentages (lots of finals appearances per tournaments), but those who have played 107 games, and seen numerous finals appearances without necessarily winning them all, ie Tatu, get proper credit as well. Also require a specific number of players for a tournament to count at all, ie 10. [insert your comments / critisisms / hate mail here] Carpe noctem. Lasombra http://www.TheLasombra.com -- Posted via Mailgate.ORG Server - http://www.Mailgate.ORG

CurtAdams

the lasombra writes: >Yes, the current rating system is flawed. >Terminally, I suspect. >Take for example the #1 player in the world for 18+ months, one of the >players from Vienna, Alex something or other. He played in 3 >tournaments, won them all, then stopped playing in tournaments for >almost 2 years. Was he really a better player than Rob Treasure all >that time? I posit that he was not. Something is pretty odd about a system in which playing 10 or so games can get you to best in the world. What if each player's score is determined by expected outcome vs. table average rather than each player? In general, I think you want to reduce the swing per game. It's basic statistics that when you have more noise in your data you must average results over a larger sample to get decent reliability. Compared to chess, you're getting noise from luck factors; this is worsened since at a table of 5 people you get 4 results but they're not independent samples; how you do against one player is correlated with how you do against another. Curt Adams (curt...@aol.com) "It is better to be wrong than to be vague" - Freeman Dyson

Anarch Troublemaker

----- Original Message ----- From: The Lasombra <thela...@hotmail.com> Newsgroups: rec.games.trading-cards.jyhad Sent: Thursday, January 24, 2002 5:14 AM Subject: Request for Comments (Doing Away With ELO Ratings) > Yes, the current rating system is flawed. > Terminally, I suspect. > > Take for example the #1 player in the world for 18+ months, one of the > players from Vienna, Alex something or other. He played in 3 > tournaments, won them all, then stopped playing in tournaments for > almost 2 years. Was he really a better player than Rob Treasure all > that time? I posit that he was not. This has nothing to do with the current rating system. if you would only count results from the last year things like this would not happen. > Is Stefan Ferenci (with 3 tournaments) really a better player than Remy > Auclair (with 14+ tournaments)? The ratings say he is. > you can bet i am :-) believe me i would love to play more tourneys but in vienna we usually have 3 sanctioned tourneys per year (max). > I would prefer to see some sort of earned runs average type of thing > myself. For example, three or four columns, all additive, no real math > for the maintainer: > > 1) Tournament Rounds played > 2) Victory Points earned (alt. Game Wins earned) > 3) Tourament Final appearances > 4) Tournament Wins > this would leave players in areas with few tourneys without a chance to have a good ranking. i would like to see sort of a winning percentage thing e.g. you played in a three round tourney going 3/2/5 (all tables were 5 player tables) so you made 10 out of 15 VP giving you a 0.66% winning percentage to honor game wins you get 1 extra point for the game win and 0,5 for a draw (these do not count to your max possible points) so in this example you would gain an extra 2.5 points so now you have 12.5 out of 15 in the finals you gain 4 VP and you won the tourney so now you have 16.5 out of 20 posiible points. to honor your victory i this tourney you get 2 extra VP, while the second gets 1 extra v.p. and the 3th 0.5 extra point. you don´t get an extra point for the table win though. so in the end you have 18.5 out of 20 points giving you 92.5% winning percentage. this system would protect players who play often (if you have 10 good tourneys and 1 bad one the bad one won´t hurt your rating that much) while not penalizing new players. it can honor game wins and tourney wins. i am still thinking how to include modfifier that deals with the rating of your opponents (a game win against rob t. and barney baker should give you more points than against some newbies). only results from the last 12 month should count. my two (european) cents. Stefan [ quoted text not captured ]

legbiter

"The Lasombra" <thela...@hotmail.com> wrote in message news:<885c9bf8d10ae618b24...@mygate.mailgate.org>... > Yes, the current rating system is flawed. > Terminally, I suspect. > > Take for example the #1 player in the world for 18+ months, one of the > players from Vienna, Alex something or other. Woisetschlager, i think. [ quoted text not captured ] Hmmm, do we really want this? i sympathise with the idea that winning a big tourney is worth more than winning a small tourney, but isn't there a danger of discouraging grass-roots play and geographically-isolated groups if we do this? How about just adding a point to every finalist for every ten whole players in the tournament? > > > [insert your comments / critisisms / hate mail here] > > > Carpe noctem. > > Lasombra > > http://www.TheLasombra.com i like your idea, but i don't like the premise which IIUC is that the ratings should tell you who is the best player. Without pitting good players against each other we will never know who is the best player, and since we have an international VTES community with tremendously-diverse local metagames, i really don't think we will EVER know who is the best player under all circumstances. Instead, i think we should look at the ratings as a method for encouraging people to play and for rewarding those who play often, and even more those who BOTH play often AND win often. The simpler such a system is, the better, i'ld have thought. On that basis, i like your system.

Andy Brown

I must say that in all honesty, no situation is either entirely fair (I could play in 20 sanctioned tourneys of only 5 people, win them all because the others are not good, and have a stonking rating) or is a headache for the organiser. I run the football tourney at work, and whilst we don't use ratings to determine position, we use pts (3 for win, 1 for draw), then goal difference, then result against other team, then matches played, it is a headache when you are trying to sort out 5 teams all on the same number of points. This is only for 11 teams. I think suggesting a new way for Todd to start positioning people, especially using multiple factors would be a nightmare. Mind you, saying that, I'm impressed by the mathematical calculation. Anyway, I think we should leave it how it is, because it will end up beingmore of a nightmare to change than leave it. Besides, is anyone really all that concerned. Unlike DCI ratings, it isn't as important to keep yours up to get invites to special tourneys. Which I must say is a good thing. Just my 10pence worth Andy Setite Ruler of Cambridge VEKN Prince -- Whether it is nobler in the mind to suffer the slings and arrows of outragous fortune, or get fat on chocolate cake, eat the cake, life is too short to worry!

JC

> Yes, the current rating system is flawed. > Terminally, I suspect. wow, quite a theatrical intro ;] > Take for example the #1 player in the world for 18+ months, one of the > players from Vienna, Alex something or other. He played in 3 > tournaments, won them all, then stopped playing in tournaments for > almost 2 years. Was he really a better player than Rob Treasure all > that time? I posit that he was not. [snip snip snip] > [insert your comments / critisisms / hate mail here] alright, i will :) (please note that i intentionally don't adress the math) here we go : i believe international rankings _make_no_sense_at_all_ even national rankings hardly make sense how can you compare a player from a small town with 10 methuselah where tournaments are held once every 6 months, with a player who lives in one of the VTES capitals of the world ? on top of the frequency of tournaments, level of play can vary quite a bit from one place to the next my personal opinion is that rankings should be city-based (for the prince to maintain) in order to be fair and meaningful in a given city, everyone can participate in tournaments held in that city if they choose to : equal opportunities if you win, you are truly better than the others in the city's ranking system as a sidenote : how did the international ranking system come about ? i'm asking because i'm wondering if it was the result of playing communities getting together and deciding they needed a wider ranking system than just that of their town, or if it was just part of a marketting policy, or something else altogether thanks for any info > Carpe noctem. a very carpe noctem to you too sir :) JC

Xian

legb...@mailandnews.com (legbiter) wrote in message news:<22fea992.0201...@posting.google.com>... [snip most of the Lasombra's post] I have to agree with Jeff that the current rating system is flawed (as an overall rating system...it seems to be a fine, internally valid indicator of recent performance). Examples probably aren't needed, but just because I can...some of the Madison guys (Cameron, Kevin, etc.) have a ranking much higher than mine, though I'm definitely in the same league as them; this is probably due to my getting hammered (just 1 VP on the day, though incredibly close to more in every game) in the last tournament I was in. Not surprisingly, this tournament was in Madison, and Kevin and Cameron did much better than I did. Presumably because of that tournament, I'm ranked a lot lower than a number of local players that I am definitely better than. And I'd compare to Josh Duffin's ranking, but I can't find it. I think he's being sneaky and listing his suburb. :) Anyway, the point is that, yup, it sure seems like the rankings really don't track anything resembling the actual rank of "who's the best V:tES player." Personally, I'm with Legbiter on what the ratings mean to me. Zilch. It's fun to look them up, but I don't play for ratings. Not that I don't play to win, but since the rating is meaningless, it doesn't matter in and of itself. On the other hand, I don't agree with Legbiter that there is no possible better rating system and that we shouldn't bother. It should be possible to come up with something better... [Jeff wrote] > > Also require a specific number of players for a tournament to count at > > all, ie 10. > > Hmmm, do we really want this? i sympathise with the idea that winning > a big tourney is worth more than winning a small tourney, but isn't Well, I'd set the minimum at 8, but that's mostly because any table with only one table just doesn't seem like an actual "tournament" to me. I've called off a tournament I was going to judge (around holiday time last year) because only 5 people showed up...it just seemed more worthwhile to play some "fun" games as opposed to wasting prize support on such a low turnout. > there a danger of discouraging grass-roots play and > geographically-isolated groups if we do this? How about just adding a > point to every finalist for every ten whole players in the tournament? I'm not sure about your point-adding idea, Legbiter. Personally, I like Josh's or Gomi's suggestion (I can't remember) of doing the average win percentage thing. It seems the most balanced suggestion I've seen. Otherwise, Curevei's suggestion of adjusting the constants for the current rating system has seemed the next most reasonable proposal. > i like your idea, but i don't like the premise which IIUC is that the > ratings should tell you who is the best player. Without pitting good IMO, that's what ratings should be for. Or at least to give you a general idea. The top 20 players in the world should be something approximating that, not just "the 20 people who have performed best at their last tournament or two." Whether or not anyone pays attention to the ratings is an entirely different matter. > players against each other we will never know who is the best player, > and since we have an international VTES community with > tremendously-diverse local metagames, i really don't think we will > EVER know who is the best player under all circumstances. Also, consider that if you got the "top 5 players" to sit down with each other, there would really be no objective way to determine who was the best player without having them play a bunch of different games with different decks and different seating arrangements. This is, of course, due to the multi-player aspect of the game, and will never really be resolved. I don't think that there's really any way to determine who "the best" is, it's just nice for a lot of people to see who's on top. > Instead, i think we should look at the ratings as a method for > encouraging people to play and for rewarding those who play often, and > even more those who BOTH play often AND win often. The simpler such a > system is, the better, i'ld have thought. On that basis, i like your > system. This all makes sense. I think any possible new ratings system ought to be less volatile and track long-term performance, rewarding performance over time rather than most recent performance. I definitely don't have the math skillz to come up with anything useful, but I'll gladly offer my input on any proposed system. :) Xian

Ben Swainbank

"The Lasombra" <thela...@hotmail.com> wrote in message news:<885c9bf8d10ae618b24...@mygate.mailgate.org>... > Yes, the current rating system is flawed. > Terminally, I suspect. > I agree and would support changing it. > > I would prefer to see some sort of earned runs average type of thing > myself. For example, three or four columns, all additive, no real math > for the maintainer: > > 1) Tournament Rounds played > 2) Victory Points earned (alt. Game Wins earned) > 3) Tourament Final appearances > 4) Tournament Wins > Ok. So then how does the ranking work? Who gets to be #1? Or is there no ranking, just a bunch of sort orders? I'd would actually like to see some sort of ranking system, even if it needn't be taken tooo seriously. -Ben Swainbank

Sten During

The Lasombra wrote: > Yes, the current rating system is flawed. > Terminally, I suspect. > Flawed? It works perfectly. In order to get an idea of a players standing I can always look up the winning decks archive - oh, sorry, that wasn't the system you're referring to :) <snip> > > I would prefer to see some sort of earned runs average type of thing > myself. For example, three or four columns, all additive, no real math > for the maintainer: > > 1) Tournament Rounds played > 2) Victory Points earned (alt. Game Wins earned) > 3) Tourament Final appearances > 4) Tournament Wins > > This both eases the burden on the maintainer, ie Todd, and gives a > better picture of how a person is performing over time. This also > allows tournament results to be entered and updated immediately, without > concern that another tournament from 1 week earlier hasn't been received > by the maintainer and having that hold up 30+ people's rankings. Easing the burden on the maintainer, while a good deed in itself, should be a matter for the choise of tools rather than something that influences the ratingsystem (if we want one at all). I'm not sure that I'd believe that an average player is a better performer than another average player because he/she participates in a lot more tournaments than the other. In this aspect ELO works very well. Actually I still think that ELO is a good way to go, but I'm not comfortable with the fact that it takes my table-ally into consideration as one opponent when I'm suddenly ousted five minutes before the table times out. LSJ had an alternative where you only count ousts, but those of you who teach tactics to any player that wants to play a rushdeck will see where that takes your predator. It might turn out that it cannot satisfy the needs for this game. If anyone is able to find any system that could be used for running a league then that system will automatically work as a valid rating as well. It will be very different, but it will work. A totally different approach is to try to measure the starting strength of each table and calculate your earned TP (tournament points) against that starting strength in comparision to your percentage of your part of that starting strength. (Table A round 1 was worth 15240 total points. Rob, entering with 3256 shared third place at the table together with Benjamin. Robs new rating is thus.....) > > Also require a specific number of players for a tournament to count at > all, ie 10. I can see why you want this, but rather than doing this I'd prefer raising the number of participants for a tournament to be official (observe that this does not mean that I want this to happen). All official tournaments ought to be official in all aspects. > > > [insert your comments / critisisms / hate mail here] > I think it's important to decide if we want a system that keep a historical track of performance or if we want one that tries to keep track of a players current capacity. The former one is probably VERY easy to implement, ie you score the inverse percentage of your position during a tournament you participate in (fifth with fiften participating gives you 66), and then just sum it all up. Sure, a mediocre performance for ten years very active gaming will place you far ahead of 100% wins for two years average gaming, but it's a system people will understand. It has nothing to do with ratings however. The latter is what we have today (I said _attempts_ to keep a current score). There are open systems as well as the closed one we use today. A closed system guarantees that the total sum of ratingpoints in the world is exactly the number of rated players times starting rating. It thus requires that you're allowed to quit while on top :) An open system is harder to keep track on, but allows that every player MUST participate in a given minumum amount of scoregenerating events during a given timeperiod in order not to see his/her score drop a percentage automatically. Strictly mathematically even a system like this is closed as it guarantees that the maximum total score equals starting score times number of participants. Sten During

Frederick Scott

The Lasombra wrote: > > Yes, the current rating system is flawed. > Terminally, I suspect. > > Take for example the #1 player in the world for 18+ months, one of the > players from Vienna, Alex something or other. He played in 3 > tournaments, won them all, then stopped playing in tournaments for > almost 2 years. Was he really a better player than Rob Treasure all > that time? I posit that he was not. > > Is Stefan Ferenci (with 3 tournaments) really a better player than Remy > Auclair (with 14+ tournaments)? The ratings say he is. ... > All the people in the top 100 that have only played in 3 tournaments > (10-12 rounds) keep their high percentages (lots of finals appearances > per tournaments), but those who have played 107 games, and seen numerous > finals appearances without necessarily winning them all, ie Tatu, get > proper credit as well. Another reason to consider changing the ELO constants. The constants control how quickly each player's number moves along the effective scale of ratings. The current numbers are extremely sensitive to a difference between a player's performance and the current ratings. That is, if a highly-rated player "loses" to a much lower-rated player (in Jyhad terms, "losing" means playing a game and getting a worse result in terms of victory points earned), then the ratings adjustments are relatively large. In Chess, where this rarely happens, such behavior is fine. In Jyhad, where this can happen a lot more frequently, it's bad - causing peoples' numbers to ricochet all over the place and making it difficult to ever accumulate a rating that goes very far from average no matter how good you might be. (If it goes beyond a certain point, you get penalized too much for a small number of losses and it brings your ratings back down towards average.) Frankly, I think the numbers chosen for the current system are not very good. (Not that I blame anyone. How can you know until you try them?) The solution is change the constants so a single result means less. The result is that it takes a player a lot longer to get to a very high rating (or a very low one) when they deserve such a rating. In a sense, this is bad since it increases the time in which ratings don't reflect actual playing skill. But for purposes of Jeff's complaint, this is actually very good. No one can play three good tournaments and rocket up to a high rating and then just stay there, laughing at the rest of the world. You'd have to play quite a few to have a rating that would compete with the better players. Now that I've really looked at the system, I'd sure hate to abandon it when it can be fixed. I see little merit to Jeff's proposal. Such statistics might make interesting "box scores" (sort of, kind of) if you like that kind of thing. But it would hardly be as interesting as a single numerical rating that actually meant something. Fred

Frederick Scott

JC wrote: > > [insert your comments / critisisms / hate mail here] > > alright, i will :) > > (please note that i intentionally don't adress the math) > > here we go : i believe international rankings _make_no_sense_at_all_ > > even national rankings hardly make sense > > how can you compare a player from a small town with 10 methuselah > where tournaments are held once every 6 months, with a player who > lives in one of the VTES capitals of the world ? Well, you do it the way it's done in the current Elo system. Since there's usually at least a little movement by individual players back and forth within different localities, the numbers can account for different skill levels. If I go to LA to play in a tournament, get squished and hosed repeatedly, and then come back and still do relatively well in Phoenix, my ratings points will decrease in LA and then I'll bring my deflated rating back home and "share the love" amongst the other local players. Thus their ratings will begin to reflect the (presumably) lesser skill level we have in our local area. As for frequency of play, this is answered by the current system being a zero-sum calculation and presumably hitting equilibrium once a player's rating begins to match his skill level. The problem with the current system is that the effective range of ratings is around 2700-3300 and the range a player's rating jumps around his skill level is something like plus or minus 100 points or more. It's hard to see much of anything in a mess like that. > my personal opinion is that rankings should be city-based (for the > prince to maintain) in order to be fair and meaningful What happens when players travel? A lot of players do. So you wind up having a bunch of listings for players who showed up and played a couple games and left. That would be a pretty ugly system, IMHO. Fred

Gomi no Sensei

In article <22fea992.0201...@posting.google.com>, legbiter <legb...@mailandnews.com> wrote: >Instead, i think we should look at the ratings as a method for >encouraging people to play and for rewarding those who play often, and >even more those who BOTH play often AND win often. The simpler such a >system is, the better, i'ld have thought. On that basis, i like your >system. Leggy, my preference is for a ratings system which rewards the players that bring the highest quality pr0n that caters best to the Head Judge's particular...um...preferences. gomi i'd judge a lot more tournaments, for sure -- Blood, guts, guns, cuts Knives, lives, wives, nuns, sluts

Frederick Scott

Sten During wrote: > > The Lasombra wrote: > > > Also require a specific number of players for a tournament to count at > > all, ie 10. > > I can see why you want this, but rather than doing this I'd prefer > raising the number of participants for a tournament to be > official (observe that this does not mean that I want this to > happen). All official tournaments ought to be official in all > aspects. Either suggestion would be very bad for people who live in areas where largish tournament are extremely difficult to hold. Believe me, when you fight attendance problems, one of the worst things you can do is to start jacking up the minimum number of player required for "official" status. Then you get this circular problem: no one's very enthusiastic about a tournament that isn't sanctioned. Or at least say, people are less enthusiastic about it. Some don't care beans about it but overall, it just doesn't help to be deprived of sanctioned status. As for Jeff's specific suggestion, I guess you'd wind up still having sanctioned tournament prize support but no ratings points would be gained or lost. That's not quite as bad. But it still doesn't help. Bottom line, if you have a really good reason to set a minimum number on a tournament then fine. But if it's just some philosophical idea about "small tournaments shouldn't count", the idea should be avoided like the plague. Remember, the success of small tournaments in an area that can only have small tournaments is absolutely crucial to getting interest in the game up to a point where you can start to have the larger tournaments. Fred

The Lasombra

"Frederick Scott" <freds64_at_...@removethis.com> wrote in message news:3C504941...@removethis.com... > > The Lasombra wrote: > > Yes, the current rating system is flawed. > > Terminally, I suspect. > Fred wrote > Now that I've really looked at the system, I'd sure hate to abandon it when > it can be fixed. I see little merit to Jeff's proposal. Such statistics > might make interesting "box scores" (sort of, kind of) if you like that kind > of thing. But it would hardly be as interesting as a single numerical > rating that actually meant something. I would also be in favor of a single number that meant something. Examine the following example: One tournament, 15 players, 3 rounds, 1 player takes 5 vp in one round with a sweep, but doesn't go to the finals. Place that sweep in the first round, and he has a high rating going into the next two rounds, where his rating is pummeled down below his starting amount. Figure that tournament the opposite way, so that the sweep against the exact same players occurs in the third round as his final performance rather than his first, and his rating starts low and will go significantly higher. Is this player's performance in the tournament (or in the game at large) different in either situatuion? No. Will his rating be different depending on which round he takes the sweep? Yes. Therefore, the ratings are worthless in their current incarnation. Getting different results from the same day of gaming has no value to me. If a system cannot be designed to calculate a rating the same way every time,(5vp rnd 1, vs 5vp rnd3), then something has to change, or you have to disregard the results. It is upon this premise that I believe the current ELO ratings are flawed and without value to this game. [ quoted text not captured ]

Frederick Scott

The Lasombra wrote: > > I would also be in favor of a single number that meant something. > > Examine the following example: > > One tournament, 15 players, 3 rounds, 1 player takes 5 vp in one > round with a sweep, but doesn't go to the finals. > > Place that sweep in the first round, and he has a high rating > going into the next two rounds, where his rating is pummeled > down below his starting amount. > > Figure that tournament the opposite way, so that the sweep against > the exact same players occurs in the third round as his final > performance rather than his first, and his rating starts low and > will go significantly higher. > > Is this player's performance in the tournament (or in the game > at large) different in either situatuion? No. > > Will his rating be different depending on which round he takes > the sweep? Yes. > > Therefore, the ratings are worthless in their current incarnation. The ratings are worthless in their current incarnation, I agree. But you're not pointing your finger at the real problem! The problem is that the rating numbers jump around so much from the results of a single game that it actually significantly *matters* in which round your hypothetical player gets his sweep. This is the kind of a game where a player can indeed get a sweep and therefore a "victory" against four other players in one round and not score another point the rest of the tournament. Yet the rating number behavior jumps significantly with each single game, as if something dramatic had been revealed by a player sweeping or getting ousted in that one game. *That's* the problem. > Getting different results from the same day of gaming has no value > to me. If a system cannot be designed to calculate a rating the > same way every time,(5vp rnd 1, vs 5vp rnd3), then something has > to change, or you have to disregard the results. Mute the ratings swings from the current game and the variation will be much smaller. And that very small variation is actually philosophically justifiable: the player's most recent performance is slightly better or worse depending on whether he swept or was ousted without a point in his "most recent game", the third game. Beyond this, worrying about such a variation is a nitpick. If I go to LA during one of those four-tournament conventions and play a tournament in the afternoon and sweep every table and then another one in the evening and get ousted without a point each game, my rating will likewise vary compared to the other way around. Yet this reflects the same result on the same day. Should it make a difference? Same issue. The bottom line is that the order things happen matters and ought to matter, if you agree that a players more recent performance is more relevant to his current rating than older history. The problem is just that right now, it matters *WAY-Y-Y-Y-Y* too much. Fred

James Coupe

In message <a2pivi$ma4$1...@panix2.panix.com>, Gomi no Sensei <go...@panix.com> writes: >Leggy, my preference is for a ratings system which rewards the players >that bring the highest quality pr0n that caters best to the Head Judge's >particular...um...preferences. Blimey. That'd save me a fortune! -- James Coupe "Never give in to them," she whispered. "No matter what PGP 0x5D623D5D EB they do or how important you feel it is to get their accep- D690ECD7A1FB457CA21 tance. Never kill part of yourself for them. Because other 3D7E668C3695D623D5D people will notice that part is missing before you do."

Robert Goudie

"CurtAdams" <curt...@aol.com> wrote in message news:20020124002343...@mb-mv.aol.com... > the lasombra writes: > > >Yes, the current rating system is flawed. > >Terminally, I suspect. > > >Take for example the #1 player in the world for 18+ months, one of the > >players from Vienna, Alex something or other. He played in 3 > >tournaments, won them all, then stopped playing in tournaments for > >almost 2 years. Was he really a better player than Rob Treasure all > >that time? I posit that he was not. > > Something is pretty odd about a system in which playing 10 or > so games can get you to best in the world. Wow. Curt Adams. How're you doin' old timer? I'd be in favor of dropping the "rank" for a player who doesn't play a minimum number of games each year. Sorta like how players who've played less than 10 games don't yet get a ranking even though their rating might be high. > What if each player's > score is determined by expected outcome vs. table average > rather than each player? In general, I think you want to > reduce the swing per game. If I remember correctly, we used to do that. It didn't reduce the swing much. I do think that reducing the swing 's a reasonable goal, though. > It's basic statistics that when > you have more noise in your data you must average results > over a larger sample to get decent reliability. I think Fred's suggestion might achieve that. It would take a while for players' ratings to move away from the middle if they only lost or gained a few points but I think we'd eventually have something worthwhile. No? > Compared to > chess, you're getting noise from luck factors; this is worsened > since at a table of 5 people you get 4 results but they're not > independent samples; how you do against one player is correlated > with how you do against another. Yep. It seems clear that an ELO system with big swings should be used solely in skill-only games. The more luck involved the less swing you should have. I look at this game much the same way I approach poker. Both are undoubtedly games of skill. But there so much luck in each hand or game that you may need to play a lot of games or hands before the better skilled player comes out on top. Our current rating system, to some degree, is similar to rating poker players after each hand. Robert Goudie Chairman, V:EKN rob...@vtesinla.org

CurtAdams

robert goudie writes: >I think Fred's suggestion might achieve that. It would take a while for players' >ratings to move away from the middle if they only lost or gained a few points >but I think we'd eventually have something worthwhile. No? Yes, I'm with Fred on this. When you have a luck factor, your "true rating" is not a number but a distribution of numbers: at any given time your official rating is a random sample from that distribution. Increased swing spreads out your equilbrium ranking distribution and increases the chance that at any given time you're below a player worse than you. But, increased swings speed approach to equilibrium. A useful experiment would be to estimate the number of tournament games a typical player plays over a career from the record. Come up with distribution of player skills and a model of how skill interacts with wins (I'm afraid that's pretty much from where the sun don't shine, although with a model you could theoretically estimate distribution from the tournament records). Then find the constants that optimize accurate player differentiation over the course of a career by brute force sims ( I really doubt you could solve it mathematically). Given this, rather than just excluding players with less than x games, you could footnote that rankings with less than x games are unreliable in that they tend to be ranked to mediocre. It's possible we may come to the conclusion that no ranking system will be particularly good, which would be useful to know if true. [ quoted text not captured ]

LSJ

Frederick Scott wrote: > Another reason to consider changing the ELO constants. The constants > control how quickly each player's number moves along the effective scale > of ratings. The current numbers are extremely sensitive to a difference The K-factor just defines the range. Any other value would have the exact same "swing" effect - the 32 is just a scale that determines the base range. The only effect changing the K factor would have is on round-off: the bigger the number, the less roundoff creeps in. -- LSJ (vte...@white-wolf.com) V:TES Net.Rep for White Wolf, Inc. Links to revised rulebook, rulings, errata, and tournament rules: http://www.white-wolf.com/vtes/

Derek Ray

In message <8GtbaFVa...@gratiano.zephyr.org.uk>, James Coupe <ja...@zephyr.org.uk> mumbled something about: >In message <a2pivi$ma4$1...@panix2.panix.com>, Gomi no Sensei ><go...@panix.com> writes: >>Leggy, my preference is for a ratings system which rewards the players >>that bring the highest quality pr0n that caters best to the Head Judge's >>particular...um...preferences. > >Blimey. That'd save me a fortune! Is goat pr0n going up in price THAT much these days? =) -- "There's no gray. There's just white that's got grubby." -- T.P.

Frederick Scott

LSJ wrote: > > Frederick Scott wrote: > > Another reason to consider changing the ELO constants. The constants > > control how quickly each player's number moves along the effective scale > > of ratings. The current numbers are extremely sensitive to a difference > > The K-factor just defines the range. Any other value would have the exact > same "swing" effect - the 32 is just a scale that determines the base > range. > > The only effect changing the K factor would have is on round-off: the > bigger the number, the less roundoff creeps in. (I'm not sure why you call it a "K-factor". I don't see that term in Appendix A. I assume you refer to the constant, 32, which is the multiplier of the difference between the score and the win probability.) I was actually proposing to change the "400" divisor in the exponent of the divisor of the probability factor. (I'll call it the J-factor, since I'm obviously not looking at the same book you're looking at.) Josh, at one time, proposed to change the K-factor. Changing either one will help the situation, I believe. I disagree with your analysis. Changing the K-factor does not change the J-factor, and that's why the numbers will start behaving differently. If you reduce the K-factor to 16 (for example), the absolute point value of sweeping three rounds of a tournament will be halved. Then, when you get into your next tournament, the difference between your rating and someone who still has the rating you used to have will get halved over the current system. Since the J-factor has not changed, this will affect the calculation of the new win probability, making it easier to continue to earn positive points and moderating the risk of losing negative points if you lose. Simply put, it will be easier to keep going in an upward direction if you're a good player. If you increase J-factor instead of reducing the K-factor, your rating will vary as much as it did before, but the win probability will still change (over the current system) in the next tournament because you're dividing the difference in ratings by a larger number. This will ultimately allow for a wider spread in peoples' ratings, even if they oscillate by the same absolute number of points as they do now. Fred

Simon Burton

In article <885c9bf8d10ae618b24...@mygate.mailgate.org>, "The Lasombra" <thela...@hotmail.com> wrote: >Yes, the current rating system is flawed. >Terminally, I suspect. > >Take for example the #1 player in the world for 18+ months, one of the >players from Vienna, Alex something or other. He played in 3 >tournaments, won them all, then stopped playing in tournaments for >almost 2 years. Was he really a better player than Rob Treasure all >that time? I posit that he was not. [snipped much of the previous post] Yes, I too think that the number of games played in overall, some sort of average should be figured into the formula of the ratings. Another thing that bothers me is that certain players may reach high ratings based upon only localised tournaments. Their ratings may well indicate they are great players but it may also be true that they are great only within their local area. The sooner we get national/continental/world championships set up worldwide the better I think with hopefully not only the opportunity being available for cross nation, or cross continent playing but that this should be highly encouraged to really see who the top players are. I know the USA and Europe has these going already. Thanks to the efforts of John Merton (Prince of Newcastle) and Salem Christ (Prince of Canberra) it looks like we should get the Australian nationals up and running this year. If we could then somehow get a world championships going where the top players could meet to contend that'd be great, and also give more accurate results on the top ranking players? (Could WW somehow afford to get these far flung players together if they can't afford it otherwise though?) Simidh, Prince of Adelaide (who has seen two local players reach the top ten and wonders if we are -really- that good here >;-> I suspect that the isolation of Adelaide from other cities is also a factor. Hopefully that will change with the nationals in place, and we can -still- retian top players :)

Sten During

Frederick Scott wrote: > Sten During wrote: > >>The Lasombra wrote: >> >> >>>Also require a specific number of players for a tournament to count at >>>all, ie 10. >>> >>I can see why you want this, but rather than doing this I'd prefer >>raising the number of participants for a tournament to be >>official (observe that this does not mean that I want this to >>happen). All official tournaments ought to be official in all >>aspects. >> > > Either suggestion would be very bad for people who live in areas where > largish tournament are extremely difficult to hold. Believe me, when > you fight attendance problems, one of the worst things you can do is I truly hope that my post above made it totally clear that I'm in total agreement with you on this point. Small tournament, large tournament, local tournament or international one, they should all, from a ratings point of view, be equally valid. Sten During

Sten During

The Lasombra wrote: [ quoted text not captured ] All you wrote is also true for a one day chess tournament, but I as a player don't feel that the ELO system is broken there. The reason for this is that the percentage difference is indeed very different. In VTES we have a scope of some 2850 - 3300. The inflation that even the zero-sum ELO is not able to handle is mainly due to new players getting their starting point, dropping in rating and getting bored. They quit playing at, let's say 2940, and never show up again. The same is true for chess, but there we have a scope of 950 - 2300. As chess is a twosided game, even during an intensive tournament I'm unable to drop more than some 150 during a day (and believe me, 6 games of tournament chess during one day is more closely to awful than simply intensive). During one VTES tournament I can unhappily see my rating drop 180 points. The amount in itself is not as important as the percentage of the total range it represents, and therein lies the problem. Sten During

Emmanuel Martin

Frederick Scott <freds64_at_...@removethis.com> wrote in message news:<3C50B87A...@removethis.com>... [ quoted text not captured ] I think another (complementary as I fully agree with you) way to enforce meaningfull ratings could be to reduce the number of wins. I mean that there would be less random if you were considered winning against someone only if you win by at least one VP against him: This means that 0,5 vp doesn't win anymore. This would sweep the "unlucky guy gets sweeped, all others go to time limit" wich can destroy a rating by giving a player four losses for what I don't see as an awfull performance. This would also prevent a bad players from getting free rating points, as the rule may sound as "No prey ousted, no victory earned" I think that these two changes (constants and at least one vp to win) could make the system definitely work better Emmanuel

Sten During

Frederick Scott wrote: > LSJ wrote: > >>Frederick Scott wrote: >> >>>Another reason to consider changing the ELO constants. The constants >>>control how quickly each player's number moves along the effective scale >>>of ratings. The current numbers are extremely sensitive to a difference >>> >>The K-factor just defines the range. Any other value would have the exact >>same "swing" effect - the 32 is just a scale that determines the base >>range. >> >>The only effect changing the K factor would have is on round-off: the >>bigger the number, the less roundoff creeps in. >> > > (I'm not sure why you call it a "K-factor". I don't see that term in > Appendix A. I assume you refer to the constant, 32, which is the multiplier > of the difference between the score and the win probability.) > > I was actually proposing to change the "400" divisor in the exponent of the > divisor of the probability factor. (I'll call it the J-factor, since I'm > obviously not looking at the same book you're looking at.) Josh, at one time, > proposed to change the K-factor. Changing either one will help the situation, > I believe. I don't remember if K or K-1 is the maximal change in rating for one win, but anyway, that's what it is. J is the difference in rating at which a win (by the lower rated player) flats out at changing the raint with K (or K-1) points. There's a lineear distribution between 0 and J in ratingdifference, so given a J = 400 the if K = 32 the following is valid: Equal rating: Win adds 16 rating points. Winner starts at 400 higher rating: Win adds 1 rating point. Winner starts at 400 lower rating: Win adds 32 rating points. Winner starts at 200 higher rating: Win adds 8 rating points. Winner starts at 200 lower rating: Win adds 24 rating points. Eventuall errors mainly due to integer rounding off, but errors in the examples above should be within 1 point of changed rating. > > I disagree with your analysis. Changing the K-factor does not change the > J-factor, and that's why the numbers will start behaving differently. If > you reduce the K-factor to 16 (for example), the absolute point value of > sweeping three rounds of a tournament will be halved. Then, when you get > into your next tournament, the difference between your rating and someone > who still has the rating you used to have will get halved over the current > system. Since the J-factor has not changed, this will affect the calculation > of the new win probability, making it easier to continue to earn positive > points and moderating the risk of losing negative points if you lose. > Simply put, it will be easier to keep going in an upward direction if you're > a good player. Given your suggestion above, within the scope of the Japanese VTES arena (if I'm correct we don't have any players there yet), we'd quickly see their national champions leveling out at some 2150 in rating, complaining about the horrid loss of some 40 points of rating during one single game, and we're back to square one. > > If you increase J-factor instead of reducing the K-factor, your rating will > vary as much as it did before, but the win probability will still change > (over the current system) in the next tournament because you're dividing the > difference in ratings by a larger number. This will ultimately allow for a > wider spread in peoples' ratings, even if they oscillate by the same absolute > number of points as they do now. If you increase J then ratings will vary less. Assuming that you set it to 1200 then you need that difference in rating between two players before the lower-rated winner will impact the rating at its maximum. Anyway, I don't think that the variation will be very great in reality as we already have to seek out the highest rated players and place them with the lowest rated players in order to see the currently used J become effective. My guess (based on no calculations at all) is that raising J a lot will see our topranked players leveling out at some 3350 or maybe a little higher. Apparently the problem we have today is widly swinging ratings within the scope of one table where the highest rated player has a less than 150 higher rating than the lowest rated player. Sten During

Joe Churchill

"The Lasombra" <thela...@hotmail.com> wrote in message news:885c9bf8d10ae618b24...@mygate.mailgate.org... > Yes, the current rating system is flawed. > Terminally, I suspect. No the system is not flawed. > Take for example the #1 player in the world for 18+ months, one of the > players from Vienna, Alex something or other. He played in 3 > tournaments, won them all, then stopped playing in tournaments for > almost 2 years. Was he really a better player than Rob Treasure all > that time? I posit that he was not. So bump up the number of required games for a true rating 20 - 40 would be a better sample IMO. Joe C. VEKN Prince of Columbia, SC www.warghoul.com [ quoted text not captured ]

andrea

Sten During <ya...@netg.se> wrote in message news:<3C51393...@netg.se>... > > If you increase J-factor instead of reducing the K-factor, your rating will > > vary as much as it did before, but the win probability will still change > > (over the current system) in the next tournament because you're dividing the > > difference in ratings by a larger number. This will ultimately allow for a > > wider spread in peoples' ratings, even if they oscillate by the same absolute > > number of points as they do now. > > > If you increase J then ratings will vary less. Assuming that you > set it to 1200 then you need that difference in rating between two > players before the lower-rated winner will impact the rating at its > maximum. Anyway, I don't think that the variation will be very great > in reality as we already have to seek out the highest rated players > and place them with the lowest rated players in order to see the > currently used J become effective. My guess (based on no calculations > at all) is that raising J a lot will see our topranked players > leveling out at some 3350 or maybe a little higher. > Apparently the problem we have today is widly swinging ratings > within the scope of one table where the highest rated player has > a less than 150 higher rating than the lowest rated player. > > > Sten During Use 1000 instead of 400 and the biggest difference is at about 300 points of difference and you'll lose 22 instead of 28 (you are higher and you lost) I think this kind of correction is hardly effective due to the "to the power" in the formula. I think the actual method should be kept It just counts your performance based on the fact that if you're good, on the average, you will outperform the others. A new player could win a tournament and have a jump up but if he is not very good he will tend to lose more than win. His next jump down will be very large. I would add two correction: Reset the score of a player that didn't play a single tournament in, say, one year. Count points win/lost after each tournment. Every round should be counted against my score at the start of the tournament. The Archon should keep note of my results and add them togheter at the end This way you correct what noted Lasombra in an another post: Losing a game and the winning another becomes equal to win and then lose. Suppose a new arena: Ten player in the contest, everyone is 3000. two rounds and the final round 1 player 1 5vp player 6 5vp round 2 player 7 5vp (Pl1 0) player 2 5VP (pl6 0 VP) (Comment on the first two rounds: well, friends, the most intercept based deck was a Malk S&B). Under the actual rule Pl2 and Pl7 went to the final round with more points (not for the tournament obviously just for the vekn score)than 1 or 6 this is not correct: same performance and, probably, same skill just wrong order. Recalculating at the end of the tournament will correct this. just my two cents (Eurocents) Andrea

Joshua Duffin

"andrea" <andrea....@infinito.it> wrote in message news:8b06c4c.02012...@posting.google.com... > Sten During <ya...@netg.se> wrote in message news:<3C51393...@netg.se>... > > > If you increase J then ratings will vary less. Assuming that you > > set it to 1200 then you need that difference in rating between two > > players before the lower-rated winner will impact the rating at its > > maximum. Anyway, I don't think that the variation will be very great > > in reality as we already have to seek out the highest rated players > > and place them with the lowest rated players in order to see the > > currently used J become effective. My guess (based on no calculations > > at all) is that raising J a lot will see our topranked players > > leveling out at some 3350 or maybe a little higher. > > Apparently the problem we have today is widly swinging ratings > > within the scope of one table where the highest rated player has > > a less than 150 higher rating than the lowest rated player. > Use 1000 instead of 400 and the biggest difference is at about 300 > points of difference and you'll lose 22 instead of 28 (you are higher > and you lost) > I think this kind of correction is hardly effective due to the "to the > power" in the formula. I'm not sure you're reading the formula correctly. (But then I'm not sure I'm reading it correctly either. Thinking about this formula makes my brain hurt.) As I read it, it's: probability of winning = 1/(1 + 10^((opponent's rating - your rating)/400)) If you're rated 300 higher than your opponent, that's: 1/(1+10^(-0.75)) = 1/(1.177828) = 0.849 and if you *do* beat this person (ie get more VPs than them in this game), you get 32*(1-0.849) = 4.83 (round to 5?) points. (If this opponent beat you, they'd get 32*(0.849) = 27 points.) If you make the 400 part of the equation much larger (say 1200), then you can win more points (and don't lose as many) even when you have a higher rating. This is what Fred Scott has been suggesting, and I think it's probably a good idea. In this example (you're rated 300 higher than your opponent), your probability of winning would be: 1/(1+10^(-0.25)) = 1/(1.56234) = 0.640 and if you beat this opponent you'd get 32*(1-0.640) = 11.52 points; if they beat you, they'd get 20.48 points. I just looked up the USCF (chess federation) formula and found that, one, they decrease the K value for players with higher ratings (it's 32 for 0-2099, 24 for 2100-2399, 16 for 2400+), so there must be some use in doing that. And two, they use the same exponentiation factor (the 400) that the VEKN is using - that determines how "probable" it is for a higher-rated player to beat a lower-rated player. So a 300-point difference for them indicates that the higher-ranked player should win 84.9% of the time, just as it does for us; a 100-point difference means the higher-ranked player should win 64% of the time. If we don't believe that a 3300-ranked player will actually beat a 3000-ranked player 85% of the time (ie score more VPs in about 6 out of 7 games), or that a 3100-ranked player will beat a 3000-ranked player 64% of the time, we should use a larger number than 400 for the ratings- difference-divisor. Josh exponentiate!

Joshua Duffin

"The Lasombra" <thela...@hotmail.com> wrote in message news:885c9bf8d10ae618b24...@mygate.mailgate.org... > I would prefer to see some sort of earned runs average type of thing > myself. For example, three or four columns, all additive, no real math > for the maintainer: > > 1) Tournament Rounds played > 2) Victory Points earned (alt. Game Wins earned) > 3) Tourament Final appearances > 4) Tournament Wins I like this concept. I'd like to add the columns "VPs per game" and "game wins per game". Also, if you want to be able to rank people with this system (and I know I do ;-) I think it'd be good to provide rankings for every column (ie you can see who's played the most tournament rounds, who's earned the most total VPs, most total game wins, highest average game wins, most total tournament wins, etc etc). And I think I'd like to still have the "current" rating as another "column", too. Preferably with one or more of the formula alterations we've been talking about. :-) Josh opinionate!

Joshua Duffin

"Xian" <xb...@qwest.net> wrote in message news:dbc1153.02012...@posting.google.com... > I have to agree with Jeff that the current rating system is flawed (as > an overall rating system...it seems to be a fine, internally valid > indicator of recent performance). Examples probably aren't needed, > but just because I can...some of the Madison guys (Cameron, Kevin, > etc.) have a ranking much higher than mine, though I'm definitely in > the same league as them; this is probably due to my getting hammered > (just 1 VP on the day, though incredibly close to more in every game) > in the last tournament I was in. Not surprisingly, this tournament > was in Madison, and Kevin and Cameron did much better than I did. > Presumably because of that tournament, I'm ranked a lot lower than a > number of local players that I am definitely better than. And I'd > compare to Josh Duffin's ranking, but I can't find it. I think he's > being sneaky and listing his suburb. :) I figured since I live in Maryland, I should list myself there. You can easily find me with a search on "Duffin". :-) According to the registry today, you're at 3034 w/22 games, ranked 166 (I won't divulge your real name here, ask Xian if you want to look him up yourself.) I'm at 3092 w/55 games, ranked 114 (Joshua Duffin, Bethesda Maryland USA). That's not a very big difference - my "win probability" against you, according to the current formula, would be 58.2%. So if I beat you in a game, I'd get 32(1-0.582) = 13 points; if you beat me in a game, you'd get 32(1-.418) = 19 points. For whatever that's worth. :-) Josh capitulate!

Xian

"Joshua Duffin" <jtdu...@yahoo.com> wrote in message news:<a2s6pf$13r24l$1...@ID-121616.news.dfncis.de>... > I figured since I live in Maryland, I should list myself there. > You can easily find me with a search on "Duffin". :-) Aha. Even though you're the Prince of D.C., huh? Heh. And I didn't see a name search thinger. But I wasn't looking too hard. > According to the registry today, you're at 3034 w/22 games, > ranked 166 (I won't divulge your real name here, ask Xian if Heh. Or you can just check the Newsletter FAQ, where Jeff lists my name...which doesn't bother me, it's just the last place I expected to see my name come up. And yeah, I got hammered after the last couple of tournaments. My rating was sky-high after GenCon, but took a steep nosedive after placing 5th in one tournament, and then getting a total of 1 VP at the next. Not that it bothers me...it's mostly just funny. The rankings (currently) seem to exist to pacify people who want some sort of definitive answer about who's better. Now, if they were changed to actually be more "meaningful" (for values of "meaningful" approximating "non-volatile historical record of performance"), then I might start worrying about it. [snip Josh's rating] > For whatever that's worth. :-) Well, I threw it in there to take a poke at you...I was going to add the tagline "Never lost to Josh..." as my signature; mostly on the theory that we haven't played (how *did* we get through GenCon without playing each other?). And then I forgot to add that sig, and I was going to reply to my own post with the sig, and then I realized that you beat me that one time here, when I was playing the 7 Raptor deck, and you were playing your Brujah Princes. So the whole point of it was lost, and I decided to just not mention it until you brought it up now. It was a sad day for everyone. :) > capitulate! Are you using Skinny Puppy lyrics or something now? Xian probably gonna use that one deck on Sunday...

Frederick Scott

Sten During wrote: > > Frederick Scott wrote: > > (I'm not sure why you call it a "K-factor". I don't see that term in > > Appendix A. I assume you refer to the constant, 32, which is the multiplier > > of the difference between the score and the win probability.) > > > > I was actually proposing to change the "400" divisor in the exponent of the > > divisor of the probability factor. (I'll call it the J-factor, since I'm > > obviously not looking at the same book you're looking at.) Josh, at one time, > > proposed to change the K-factor. Changing either one will help the situation, > > I believe. > > I don't remember if K or K-1 is the maximal change in rating for one > win, but anyway, that's what it is. It is whatever '32' is in the equation defined in Appendix A, no more, no less. I suspect that K-1 is effectively the maximum change in rating for one win, but to define as only that ignores that constant's effect on the entire process. > J is the difference in rating at which a win (by the lower rated player) > flats out at changing the raint with K (or K-1) points. No. J helps define the _rate_ by which the exponential curve flattens out but there is no strict point at which the curve becomes "flat". You oversimplify far too much. > There's a > lineear distribution between 0 and J in ratingdifference, so given a > J = 400 the if K = 32 the following is valid: > Equal rating: Win adds 16 rating points. > Winner starts at 400 higher rating: Win adds 1 rating point. > Winner starts at 400 lower rating: Win adds 32 rating points. Incorrect. (Assuming I read the appendix correctly.) The win probability for two players with a 400 difference in their ratings are: 1/(10**(400/400)+1) = 1/(10**1+1) = 1/(10+1)= 1/11, if the lower guy wins 1/(10**(-400/400)+1)= 1/(10**-1+1)=1/(0.1+1) = 1/1.1 = 10/11 otherwise That means the win that if the winner starts at a 400 higher rating, the points shifted = 32 * (1 - 10/11) = 32 * (1/11) or about 3 points. If the winner is the lower rated guy the equation is 32 * (1 - 1/11) = 32 * (10/11) or about 29 points. > Winner starts at 200 higher rating: Win adds 8 rating points. > Winner starts at 200 lower rating: Win adds 24 rating points. Again, let's do the actual math. 1/(10**(200/400)+1) = 1/(10**1/2+1)=1/(3.162278+1) = 1/4.162278 = 0.240253 1/(10**(-200/400)+1) = 1/(10**-1/2+1)=1/(0.316228+1) = 1/1.316228 = 0.759747 So the actual points are 8 and 32 as you say. Don't know if did these correctly or just got the right answers by accident. > > I disagree with your analysis. Changing the K-factor does not change the > > J-factor, and that's why the numbers will start behaving differently. If > > you reduce the K-factor to 16 (for example), the absolute point value of > > sweeping three rounds of a tournament will be halved. Then, when you get > > into your next tournament, the difference between your rating and someone > > who still has the rating you used to have will get halved over the current > > system. Since the J-factor has not changed, this will affect the calculation > > of the new win probability, making it easier to continue to earn positive > > points and moderating the risk of losing negative points if you lose. > > Simply put, it will be easier to keep going in an upward direction if you're > > a good player. > > Given your suggestion above, within the scope of the Japanese VTES arena > (if I'm correct we don't have any players there yet), we'd quickly see > their national champions leveling out at some 2150 in rating, > complaining about the horrid loss of some 40 points of rating during > one single game, and we're back to square one. I have no idea what any of this means. What is the "Japanese VTES area", why is it significant, why would their champions level off at a rating 850 points below the starting level, how does the number '40' enter into all this, and why would it be a bad number???? For one thing, the number of points that one can lose in a single game is dependent on whether it's a 4-player or 5-player game. I confess, you've absolutely left me the dust with what appears to be a mind-boggling array of successive non sequiturs. > > If you increase J-factor instead of reducing the K-factor, your rating will > > vary as much as it did before, but the win probability will still change > > (over the current system) in the next tournament because you're dividing the > > difference in ratings by a larger number. This will ultimately allow for a > > wider spread in peoples' ratings, even if they oscillate by the same absolute > > number of points as they do now. > > If you increase J then ratings will vary less. Assuming that you > set it to 1200 then you need that difference in rating between two > players before the lower-rated winner will impact the rating at its > maximum. There is no such maximum as far as I can see. It appears to me that you and I are looking at a different set of equations. I would suggest you go to Appendix A in the tournament rules and look at what's written there and go from there. To me, it sure looks like if you reset J to 1200, ratings will vary more. It will give superior players a chance to go higher and stay higher since the difference in ratings between them and the lower-rated players they oppose will affect the ratings changes _less_. Once they reach their approximately correct rating, they'll still move up and down numerically as much they do now, but it will move up and down around a number that's a lot further away from the average players' ratings. Fred

Frederick Scott

andrea wrote: (I wrote:) > > > If you increase J-factor instead of reducing the K-factor, your rating will > > > vary as much as it did before, but the win probability will still change > > > (over the current system) in the next tournament because you're dividing the > > > difference in ratings by a larger number. This will ultimately allow for a > > > wider spread in peoples' ratings, even if they oscillate by the same absolute > > > number of points as they do now. > > Use 1000 instead of 400 and the biggest difference is at about 300 > points of difference and you'll lose 22 instead of 28 (you are higher > and you lost) > I think this kind of correction is hardly effective due to the "to the > power" in the formula. I'm not sure where you got your numbers but I think it has to be effective. In the current system: Difference = 400, lose 29 points Difference = 200, lose 24 points Difference = 0, lose 16 points Difference = -200, lose 8 points Difference = -400, lose 3 points. That's a pretty extreme curve. Change J to 1000 and this same curve is spread out to 1000, 500, 0, -500, and -1000 respectively. I can't see how that _wouldn't_ (double-negative intended) be very effective in changing how the rating system behaves. Fred

Sten During

Frederick Scott wrote: > Sten During wrote: <snipping a lot where we obviously are in agreement at large> >>> >>Given your suggestion above, within the scope of the Japanese VTES arena >>(if I'm correct we don't have any players there yet), we'd quickly see >>their national champions leveling out at some 2150 in rating, >>complaining about the horrid loss of some 40 points of rating during >>one single game, and we're back to square one. >> > > I have no idea what any of this means. What is the "Japanese VTES area", > why is it significant, why would their champions level off at a rating 850 > points below the starting level, how does the number '40' enter into all > this, and why would it be a bad number???? For one thing, the number of > points that one can lose in a single game is dependent on whether it's a > 4-player or 5-player game. I confess, you've absolutely left me the dust > with what appears to be a mind-boggling array of successive non sequiturs. Here's what I get when I'm not checking what I write ;) Leveling out at 3150 of course, not 2150 as I wrongly wrote. What I was aiming at was that within the scope of a 'new' region with a lot of players starting at 3000 rating, if you switch from 32 to 16, then the percentage swing during one game will not vary, so with the hypothectical "Japanese VTES arena", then instead of top-ranked players leveling out at 3300 and complaining about some 80 loss during one round, you'd have top-ranked players levelling out at 3150 complaining about drops of 40, which from a percentage point of view is equally bad. > > >>>If you increase J-factor instead of reducing the K-factor, your rating will >>>vary as much as it did before, but the win probability will still change >>>(over the current system) in the next tournament because you're dividing the >>>difference in ratings by a larger number. This will ultimately allow for a >>>wider spread in peoples' ratings, even if they oscillate by the same absolute >>>number of points as they do now. >>> >>If you increase J then ratings will vary less. Assuming that you >>set it to 1200 then you need that difference in rating between two >>players before the lower-rated winner will impact the rating at its >>maximum. >> > > There is no such maximum as far as I can see. It appears to me that you and > I are looking at a different set of equations. I would suggest you go to > Appendix A in the tournament rules and look at what's written there and > go from there. To me, it sure looks like if you reset J to 1200, ratings > will vary more. It will give superior players a chance to go higher and > stay higher since the difference in ratings between them and the lower-rated > players they oppose will affect the ratings changes _less_. Once they reach > their approximately correct rating, they'll still move up and down numerically > as much they do now, but it will move up and down around a number that's a > lot further away from the average players' ratings. > > Fred > No, will vary less, by which I mean that high-ranking players will drop less during the course of one round. Bah, why are we arguing over a point where we apparently pretty much agree :) I do agree with you (otherwise would be mathematically totally inconsistent), that we'd see our topranked players levelling out at higher ratingvalues than today if you change J from 400 to 1200, but as our topranked players already level out at less than todays J-value I doubt that the 3300 will rise very much, ie the effect won't be as great as I believe that you think. I don't recall if chess has a J-value of 400 or 600, but topranked players have some 2300 rating (starting at 1300), which is a higher rise compared to the J-value than we have with VTES. Admittedly there's less room for 'bad luck' when you play chess ;) Sten During

Joshua Duffin

"Xian" <xb...@qwest.net> wrote in message news:dbc1153.02012...@posting.google.com... > Aha. Even though you're the Prince of D.C., huh? Heh. And I didn't > see a name search thinger. But I wasn't looking too hard. The area search thing works for names as well. Better for names than areas in some cases - if you want North Carolina you can get all the North states with "North" or both Carolinas with "Carolina", but not NC by itself. Happily Minnesota doesn't have that problem. Or Maryland. But for DC you want to search on District rather than Columbia. :-) > And yeah, I got hammered after the last couple of tournaments. My > rating was sky-high after GenCon, but took a steep nosedive after > placing 5th in one tournament, and then getting a total of 1 VP at the > next. Not that it bothers me...it's mostly just funny. The rankings > (currently) seem to exist to pacify people who want some sort of > definitive answer about who's better. Now, if they were changed to > actually be more "meaningful" (for values of "meaningful" > approximating "non-volatile historical record of performance"), then I > might start worrying about it. Yeah... I'd like it to be less volatile too. Maybe Fred Scott's suggestion will get some play. (I think he's right that increasing the "J" value (the 400) is the way to go, though I'd guess that it should be at *least* 1200; I don't want to assume that a player at 3400 is more than 68% likely to beat a player at 3000 (what 1200 would do) and 60% (this would be about 2300 for "J") would be better.) (Also, Sten and others, note that reducing "K" (the 32) *does* have an indirect effect on effective rating range - if people's ratings don't start out by getting very far away from 3000, they won't start getting "reduced awards for having a high rating" until a later point.) > Well, I threw it in there to take a poke at you...I was going to add > the tagline "Never lost to Josh..." as my signature; mostly on the > theory that we haven't played (how *did* we get through GenCon without > playing each other?). Ahhh, now it all makes sense. I'm sure it was easy to get through GenCon w/o playing in the same 'official' game - there were 75 players in that final qualifier and I didn't play in the "losers" event. Why we didn't play in any of the same pickup games is a good question though... guess we were just unlucky. Maybe we'll do better this year. :-) > And then I forgot to add that sig, and I was going to reply to my own > post with the sig, and then I realized that you beat me that one time > here, when I was playing the 7 Raptor deck, and you were playing your > Brujah Princes. So the whole point of it was lost, and I decided to > just not mention it until you brought it up now. It was a sad day for > everyone. :) Heh! That game was so long ago I don't even remember what happened in it. And wasn't it like an eight-player game or something? > > capitulate! > > Are you using Skinny Puppy lyrics or something now? No, I picked some -ate word for some reason and then it seemed like a good idea to keep using them. What, am I supposed to make sense or something? > probably gonna use that one deck on Sunday... Better than using that no deck. :-) Josh used that blood bros deck on saturday (and, surprisingly, won)... report coming tomorrow.

Frederick Scott

Joshua Duffin wrote: > Yeah... I'd like it to be less volatile too. Maybe Fred > Scott's suggestion will get some play. (I think he's right > that increasing the "J" value (the 400) is the way to go, > though I'd guess that it should be at *least* 1200; I don't > want to assume that a player at 3400 is more than 68% likely > to beat a player at 3000 (what 1200 would do) and 60% (this > would be about 2300 for "J") would be better.) I have to say, after I posted the "1000" suggestion, I thought it over and began to wonder if it would hurt anything to make it 1500 or even 2000. I'm pretty sure the effective spread on these things is going to be no more than +/- 50% to 60% of "J" if people can play enough games to get up there. J=2000 means that once you get up to 4000, you are gaining 8 pts for a win and losing 24 pts. for a loss against 3000-rated players. You'd have to pretty good to maintain that for long, I think. And it would take you about 60-70 straight points or 15-18 5-player table sweeps to get you there. Assuming you can win three contests for every one you lose (equilibrium for a 4000-rating in a J=2000 system), it appears that you should be able to get there in 30-35 5-player games or 40-45 4-player games. (Assuming you constantly play against 3000-level opposition.) (By the way, I'm kicking this all out off the top of my head. Anyone wants to check my numbers rigorously, be my guest.) Fred

Xian

"Joshua Duffin" <jtdu...@yahoo.com> wrote in message news:a34i8f$15h848$1...@ID-121616.news.dfncis.de... [rating = "non-volatile historical record of performance"?] > Yeah... I'd like it to be less volatile too. Maybe Fred > Scott's suggestion will get some play. (I think he's right It seems like a good idea to me too, but then again, that might be an argument for the opposition. :) > Why we didn't play in any of the same pickup games is a > good question though... guess we were just unlucky. Maybe > we'll do better this year. :-) Yeah, that's what confused me. Even the last night, we were talking about having to play a game, and then it never occurred. Heh. [7 Raptor game] > Heh! That game was so long ago I don't even remember what > happened in it. You let me build up because you wanted to see what would happen, then decided I was probably going to get annoying if I got more than 4 Raptors, rushed, knocked my Raptor guy into torpor by hanging onto a Torn Signpost & a Blur (or something like that), and then proceeded to oust me. Then you got ganged up on because you were an obviously good player (out of the 8 at the table), and because there was a lot of contention about Tor voting for your Brujah Justicar in return for a Consanguineous Boon. Oh yeah, and he was my prey. Or something. And then Tor probably would have taken the rest of the table with my Ventrue vote & bloat deck, but you guys had to leave because you had some sort of "family committment". I wouldn't remember this all that well, except I was paging through old posts fairly recently. And yes, it was an obscene 8-player game. Happily, we rarely see those anymore. And when someone suggests one, I sit out and annoy them. Except now maybe I'll play with my Ventrue voters Con Ag/Slave Auction deck. Hehe. Just to be a bastard. > No, I picked some -ate word for some reason and then > it seemed like a good idea to keep using them. What, am > I supposed to make sense or something? Not necessarily. After further thought, "capitulate" sounded more like Front Line Assembly or Clock DVA than Skinny Puppy. Xian has only seen one 8-player game in the last 6 months...

Joshua Duffin

"Xian" <xi...@waste.org> wrote in message news:a34nt0$15etvr$1...@ID-123937.news.dfncis.de... > > [7 Raptor game] > You let me build up because you wanted to see what would happen, then > decided I was probably going to get annoying if I got more than 4 Raptors, > rushed, knocked my Raptor guy into torpor by hanging onto a Torn Signpost & > a Blur (or something like that), and then proceeded to oust me. Then you > got ganged up on because you were an obviously good player (out of the 8 at > the table), and because there was a lot of contention about Tor voting for > your Brujah Justicar in return for a Consanguineous Boon. Oh yeah, and he > was my prey. Or something. And then Tor probably would have taken the rest > of the table with my Ventrue vote & bloat deck, but you guys had to leave > because you had some sort of "family committment". Yeah, dinner or something, I dunno. > I wouldn't remember this all that well, except I was paging through old > posts fairly recently. That'd do it. heh! Funny game. Raptors, ha! Eight players, ha! ;-) > And yes, it was an obscene 8-player game. Happily, we rarely see those > anymore. And when someone suggests one, I sit out and annoy them. Except > now maybe I'll play with my Ventrue voters Con Ag/Slave Auction deck. Hehe. > Just to be a bastard. Yeah, the "number of players" cards are definitely the ones to use in that kind of game, which is (besides the absurd time between turns) the main problem with them in my mind. Parity Shift's even worse... :-) Josh recapitulate...

Frederick Scott

Sten During wrote: > > Frederick Scott wrote: > > > Sten During wrote: > Here's what I get when I'm not checking what I write ;) > Leveling out at 3150 of course, not 2150 as I wrongly wrote. > > What I was aiming at was that within the scope of a 'new' region with > a lot of players starting at 3000 rating, if you switch from 32 to > 16, then the percentage swing during one game will not vary, so with > the hypothectical "Japanese VTES arena", then instead of top-ranked > players leveling out at 3300 and complaining about some 80 loss during > one round, you'd have top-ranked players levelling out at 3150 > complaining about drops of 40, which from a percentage point of view is > equally bad. Top-ranked players would still level out at 3300 if they did before***. You haven't changed the J-factor (referred to in other posts), thus the "thing" that makes them level out is still the same. The K-factor does not control what creates the equilibrium, only how rapidly the players arrive at it. Remember, what's creating the equilibrium is the points that get calculated when a player wins or loses. Win and you get X points given your current rating and the current rating of the opposition. Lose and you lose Y points. When you so far above 3000, the ratio of X to Y starts to get sufficiently small that it counteracts your ability to win more often than you lose. Do remember: it's the *RATIO* of these two numbers that creates the equilibrium, not their raw magnitude. J and K both factor into their raw magnitude, but only J and the difference between the rating factors affect the ratio. Since the former has not changed and the latter is self-correcting, essentially neither of these has changed. The ratio will still be the same no matter what you do to K. The problem, if looked at in terms of K, is that K is so relatively large right now it causes oscillation so great that a player's effective range blurs with any other player who can keep his rating over 3000 most of the time. You wind up being unable to tell who's really good and who's just had a recent run of luck. *** - I don't actually believe anyone levels at 3300 now. If they did, you'd see those guys running up over 3400 from time to time. I think the ratings that go up over 3300 are the best players on a recent upswing, before they run themselves back down to 3100. 3200 is probably about the best average anyone maintains over time right now. > >>If you increase J then ratings will vary less. Assuming that you > >>set it to 1200 then you need that difference in rating between two > >>players before the lower-rated winner will impact the rating at its > >>maximum. > > > > There is no such maximum as far as I can see. It appears to me that you and > > I are looking at a different set of equations. I would suggest you go to > > Appendix A in the tournament rules and look at what's written there and > > go from there. To me, it sure looks like if you reset J to 1200, ratings > > will vary more. It will give superior players a chance to go higher and > > stay higher since the difference in ratings between them and the lower-rated > > players they oppose will affect the ratings changes _less_. Once they reach > > their approximately correct rating, they'll still move up and down numerically > > as much they do now, but it will move up and down around a number that's a > > lot further away from the average players' ratings. > > No, will vary less, by which I mean that high-ranking players will drop > less during the course of one round. Bah, why are we arguing over a > point where we apparently pretty much agree :) We do not agree here. You are incorrect: increase J and high-ranking players will lose less points if they lose but gain more points if they win. This does not sound like "varying less" to me. It sounds like being relatively rewarded more or punished less for a result given the same point differential with a given opponent. You have to be careful about your terminology, I guess. > I do agree with you (otherwise would be mathematically totally > inconsistent), that we'd see our topranked players levelling out at > higher ratingvalues than today if you change J from 400 to 1200, but as > our topranked players already level out at less than todays J-value > I doubt that the 3300 will rise very much, ie the effect won't be as > great as I believe that you think. The effect should be directly proportional to the increase in J-value. If you quadruple J, then the effective range of ratings should quadruple, given sufficient games played to achieve true equilibrium. That's "as much as I think". If there's any flaw in the reasoning, I'd like to hear it. Fred

Sten During

Frederick Scott wrote: >>>Sten During wrote: >>> >>complaining about drops of 40, which from a percentage point of view is >>equally bad. >> > > Top-ranked players would still level out at 3300 if they did before***. > You haven't changed the J-factor (referred to in other posts), thus the > "thing" that makes them level out is still the same. The K-factor does > not control what creates the equilibrium, only how rapidly the players > arrive at it. > > Remember, what's creating the equilibrium is the points that get calculated > when a player wins or loses. Win and you get X points given your current > rating and the current rating of the opposition. Lose and you lose Y points. > When you so far above 3000, the ratio of X to Y starts to get sufficiently Look at it this way. The number of games 'lost' and 'won' for any one individual player will not vary dependent on a ratingsystem. What happens when you cut K in half without changing J is that you'll have a somewhat less punishing system for players with rising ratings. I'm still referring to a 'new' region, so all players start at 3000 rating. A player that with the current numbers have risen to 3200 should with K/2 as new K probably be found in the 3120 - 3130 range, which admittedly is higher from a percentage point of view, but sooner or later you'll find some kind of equilibrum (unsure spelling), and if we today find it at 3300 (the rating from which any topranked player will drop uncontrollably) then with K/2 we'll see it at 3160 - 3180, and the complaints will be pretty much the same. If you go for K/4 or K/8 then it'll be another story because the rounding off effects will take over in combination with standard deviation being applied on a much smaller span of possibilities. > > *** - I don't actually believe anyone levels at 3300 now. If they did, you'd > see those guys running up over 3400 from time to time. I think the ratings > that go up over 3300 are the best players on a recent upswing, before they > run themselves back down to 3100. 3200 is probably about the best average > anyone maintains over time right now. I actually don't think that anyone levels at all now, which is the core problem with todays implementation. Players just swing up and down almost uncontrollably. >>> >>No, will vary less, by which I mean that high-ranking players will drop >>less during the course of one round. Bah, why are we arguing over a >>point where we apparently pretty much agree :) >> > > We do not agree here. You are incorrect: increase J and high-ranking players > will lose less points if they lose but gain more points if they win. This does > not sound like "varying less" to me. It sounds like being relatively rewarded > more or punished less for a result given the same point differential with a > given opponent. You have to be careful about your terminology, I guess. Ok, we use terminology differently, but both paragraphs above say the same thing. Less punishment and more gain. ELO implicitly forces more gain in combination with less punishment. > > >>I do agree with you (otherwise would be mathematically totally >>inconsistent), that we'd see our topranked players levelling out at >>higher ratingvalues than today if you change J from 400 to 1200, but as >>our topranked players already level out at less than todays J-value >>I doubt that the 3300 will rise very much, ie the effect won't be as >>great as I believe that you think. >> > > The effect should be directly proportional to the increase in J-value. If > you quadruple J, then the effective range of ratings should quadruple, > given sufficient games played to achieve true equilibrium. That's "as much > as I think". If there's any flaw in the reasoning, I'd like to hear it. > > Fred > That would have been true if players actually did manage to gain more rating than the J-value, which they don't. Switching J from 400 to 1200 will not make us see players make some 3900 rating (provided that todays max actually is 3300). I don't know the 'new' rating that would become the max, but a guess is somewhere in the 3500 - 3600 range. Now, my differing in opinion has not been based on by which metric ELO is or is not flawed. I don't see ELO as flawed per se, but rather that todays definition on a 'win' differs from the new tournament rules which since this year promotes an 'all or nothing' style of gaming compared to the old where you should gather a maximum sum of VP during the course of three rounds. Thus I would prefer a rating system that somehow promoted Table Wins a bit extra or even one that only counts Table Wins compared to what we have today, and this opinion of mine is only based on the new rules that explicitly states that it's better to get one Table Win and get ousted first during the other two rounds than grabbing 2 VP each round unless one of those rounds also gives you a Table Win. Whatever is primarily most important to get you to the finals should also be most important to count when we attempt to measure who, in the long run, is the better player. If we don't do this, then we'll stay with the only actually accepted 'ratingsystem' we have today which is based on 'I have a feeling that this player usually makes it to the finals more often than that player, so this player is better'. Sten During

Frederick Scott

[ quoted text not captured ] I already explained why this wasn't true: the conditions which cause the equilibrium (how good the player is and what the ratio of points he gets for winning vs. losing) don't change when you change only K and not J. Therefore, the equilibrium MUST be the same as it is now. I don't know how you're fooling yourself but your conclusions are incorrect. > > *** - I don't actually believe anyone levels at 3300 now. If they did, you'd > > see those guys running up over 3400 from time to time. I think the ratings > > that go up over 3300 are the best players on a recent upswing, before they > > run themselves back down to 3100. 3200 is probably about the best average > > anyone maintains over time right now. > > I actually don't think that anyone levels at all now, Sure they do. The just swing up and down in a range of 200 - 250 points. Maybe more for a particularly good or bad swing. But if they didn't level, then you'd see players wandering around to any level, given sufficient games - even over 4000, etc. > >>I do agree with you (otherwise would be mathematically totally > >>inconsistent), that we'd see our topranked players levelling out at > >>higher ratingvalues than today if you change J from 400 to 1200, but as > >>our topranked players already level out at less than todays J-value > >>I doubt that the 3300 will rise very much, ie the effect won't be as > >>great as I believe that you think. > > > > The effect should be directly proportional to the increase in J-value. If > > you quadruple J, then the effective range of ratings should quadruple, > > given sufficient games played to achieve true equilibrium. That's "as great > > as I think". If there's any flaw in the reasoning, I'd like to hear it. > > That would have been true if players actually did manage to gain more > rating than the J-value, which they don't. > Switching J from 400 to 1200 will not make us see players make some > 3900 rating (provided that todays max actually is 3300). I don't know > the 'new' rating that would become the max, but a guess is somewhere > in the 3500 - 3600 range. All I can see is that these conclusions are incorrect - and you made a sloppy error in extending my reasoning: I did not say say that switching J from 400 to 1200 will result in players having 3900 ratings. 3300 is a result of the best player's running their ratings up to the high point of the range in ratings: 3200 effective rating + 100 pts on the high end of the +/- 100 points of their effective range. Triple J and you triple how far their range is from the middle. Three times 200 is 600, therefore you will see player's ratings go from 3600 +/- 100 points or up to 3700. (Assuming I'm correct that the best players are *averaging* 3200 points. Since we don't know that for sure, it's hard to tell exactly what will happen.) The rest of your conclusions are unsupported and, I believe, simply incorrect. Multiply J by some number and the *center* of the highest effective range will move away from 3000 by that same multiple. Fred