An extra day, one that only exists every four years. And I am spending this one doing homework. Good thing I love statistics or I would be slightly put out. Last week, before I realized how much work I actually needed to do today, I was planning a leap-day party full of chocolate strawberries and watching "Leap Year" with my roommates. Therefore, I somewhat follow the Rules of Leisure according to Kate Fox. True, I am not improving my apartment with a DIY, but my leisure party follows the privacy so desired by the English.
After reading Rules of Play, I think I will mostly like "playing" in England, but only after a hard day's work of statistics of course. I will admit, I do not think I will ever get drunk and fight someone, so I'll have to just observe that form of social interaction. With the private and domestic pursuits, my favorite was "visiting grand country houses." If it is indeed "one of the most popular national pastimes," I will apparently need to make visiting grand country houses a priority in order to fully submerse myself in the English culture. At least now I will not feel quite as guilty for visiting as many as I am planning for the field study. However, I was surprised that such country houses were for the English themselves and not simply for the tourists. Not only do the English visit the houses, they enjoy it. But I won't argue; I think it is a splendid idea. On my study abroad, girls would complain about all the houses we were forced to visit. The complaining never made sense to me. To me, the houses represented a dying way of life, a life of luxury that most will never have. And I love architecture, and some of those English houses are perfect.
Another pastime I am anxious to implement is "the pub."My first time in London, I did not actually visit any pubs. Legally, I was old enough so that cannot be my excuse. Maybe I was too busy visiting grand houses to visit the pubs. Whatever the reason, now that I am armed with the rules of the pub, I am ready to visit as many as possible. Mostly, I think it will be a great opportunity to see the workings of English culture.
Tuesday, February 28, 2012
Source (2/29)
Marley, C.J. and D.C. Woods. "A comparison of design and model selection methods for supersaturated designs." Computational Statistics and Data Analysis. 54.12 (2010): 3158-3167. Electronic.
Before developing and running a full-blown experiment, a screening experiment is run to discover active factors so statisticians know what to include in the real experiment. A supersaturated design, "in which the number of factors exceeds the number of runs," is used when a large experiment is impractical. However, there is not enough information about the comparison and evaluation of various methods for supersaturated designs. This paper utilizes simulations using different sample sizes and number of active factors.
In the paper, it was stated that the "most flexible design construction methods are algorithmic." In other words, simulations studies reign supreme in supersaturated designs. Another helpful aspect of this article was the list of analysis of data, how the simulations were run to obtain the appropriate results.
Before developing and running a full-blown experiment, a screening experiment is run to discover active factors so statisticians know what to include in the real experiment. A supersaturated design, "in which the number of factors exceeds the number of runs," is used when a large experiment is impractical. However, there is not enough information about the comparison and evaluation of various methods for supersaturated designs. This paper utilizes simulations using different sample sizes and number of active factors.
In the paper, it was stated that the "most flexible design construction methods are algorithmic." In other words, simulations studies reign supreme in supersaturated designs. Another helpful aspect of this article was the list of analysis of data, how the simulations were run to obtain the appropriate results.
Monday, February 27, 2012
Source (2/27)
Biedermann, S. and Woods, D.C. "Optimal designs for generalized nonlinear models with application to second harmonic generation experiments." Journal of the Royal Statistical Society, Series C. 60.2 (2011): 281-299. Electronic.
The paper extends the theoretical basis for non-linear regression using Bayesian design to cluster design. For experiments where the errors are not believed to be normally distributed, non-linear parametric regression models are needed to "describe the influence of one or more explanatory variables on a response." More specifically, "Generalized non-linear models extend non-linear regression models to allow non-normally distributed error structures." After formulating the procedure, the authors use various tests to determine the robustness of GNM (generalized non-linear models).
This article outlines the mathematical and practical tests statisticians use in determining properties of new theories. If my project tends to the more theoretical side, this provides a wonderful basis for usual tests.
The paper extends the theoretical basis for non-linear regression using Bayesian design to cluster design. For experiments where the errors are not believed to be normally distributed, non-linear parametric regression models are needed to "describe the influence of one or more explanatory variables on a response." More specifically, "Generalized non-linear models extend non-linear regression models to allow non-normally distributed error structures." After formulating the procedure, the authors use various tests to determine the robustness of GNM (generalized non-linear models).
This article outlines the mathematical and practical tests statisticians use in determining properties of new theories. If my project tends to the more theoretical side, this provides a wonderful basis for usual tests.
Saturday, February 25, 2012
Splines in Linear Regression (LJ 2/27)
I worry about you Averyl, having to read some of my learning journals. This one will be brutal.
A lot of the papers I have read (skimmed) written by Dr. Woods involve modern regression model methods, such as splines (surprisingly, that graph is an example of a linear, truly linear, model using splines). Currently, I am taking a class on modern regression model (although in my case, modern means 1980, so not cutting edge methods like Dr. Woods), and I realized I will definitely need to understand this core material for my project. In order to make sense of what I have learned, I will briefly detail some modern regression methods and how they will relate to my project.
Data is not linear, regardless of how much statisticians wish it were. Yet there is some much clean and intuitive theory about linear models. Mathematically, linearity is optimal. Instead of wading through murky mathematics and kooky calculations (many of which have been proved impossible to solve), statisticians have adapted linear ideas to fit curvy data. Some methods include splines, smoothing kernels, transformations, and automatic smoothers. Current thought on experimental design relies heavily on such methods, especially splines and smoothers.
Splines: Sometimes, between different experimental groups, the effects of a drug is are different, i.e. the slopes are different. Therefore, one needs to use different lines to predict for the different groups. But you run into problems with continuity, so you extend the basis. (NOTE to self: review linear algebra; Dr. Woods is very mathematical.) When designing experiments, it is important to take the supposed differences into account in order to reduce both variance and bias, the paradigm of statistics.
Smoothers: When the underlying distribution of the points is not a linear, or even a piecewise collection of them, you can only predict point by point. Using the same knot selection idea as in splines, an alternative approach is to put knots at all distinct x values and control the fit of the line (actual linear line) through regularization. Using a smoother matrix formulated from the data, the effective degrees of freedom can be chosen. In design theory and Dr. Woods papers, there is some comment about the selection of effective degrees of freedom. In class, we use a greedy algorithm to chose them, but I would like to be able to learn more about other efficient and conservative ways in choosing the effective degrees of freedom.
Such methods are very applicable in design theory, or at least the more sophisticated and elegant relatives of the above methods. It is important for me to have a strong basis in the above methods so I may have somewhere to build from.
Thursday, February 23, 2012
Modeling My Personality (LJ 2/24)
Taking the test, I realized there were many principles I wish I could answer "no" to but honor-bounded to the truth, I had to answer "yes." For example, "When solving a problem, you would rather follow a familiar approach than seek a new one. " I study statistics; I should be actively seeking new approaches to solving problems. New ideas and approaches are what drives statistics forward. But I love knowing one approach will work. On the flip side, there were questions I had to answer "no," with "You have good control over your desires and temptations." There is a scrumptious, delicious Ghirardelli brownie in my pantry right now, and I do not think I can hold out much longer.
Enough confessions. My profile from two different tests was surprisingly similar. One stated I was
- Very expressed introvert (surprise!)
- Slightly expressed sensing
- Slightly expressed thinking
- Slightly expressed judging
- 71.43% introverted
- 52.94% sensing
- 66.67% thinking
- 76.92% judging
Though not a deal breaker, being an introvert will complicate my project a wee bit. Fortunately, I am not formally interviewing people, but I still will need to talk with people as I do my participating-observing. However, I will be solely in an academic setting, i .e. working with very few people at a time. Therefore, establishing relationships with my fellow studiers will be slow but doable. Also, being an introvert will be helpful in the statistics half of my project. I will have to work alone (excluding my computer) for hours. In fact, on the careers that matched my profile, statistician was number four. According to myersbriggs.org, I have the capability to "decide logically what should be done and work toward it steadily, regardless of distractions." Sounds brilliant for working on statistics.
Sensing/ Intuition (because I basically was fifty-fifty, and I liked the description of both)
The combination of these two personality traits is basically the definition of statistics. Statisticians focus on the fundamental information and then they interpret the results to add meaning to numbers. In regards to the communication portion, it will be helpful to have the basic definitions and theories of communication and then apply and observe (or not) them in the academic world of establishing collaboration.
Thinking
Of course I think: "logic and consistency" is all I am about. Not really one for special cases; they always make proofs and theorems ugly and messy. So again, this is advantageous for the statistics portion of my project. As horrible as it is to admit, academic relations are not really about genuine liking, though it does help. Instead, they are formed for the purpose of academic collaboration, so it is helpful to be thorough and dependable in collaborations. My tendency to look at facts first not people may be a stumbling-block in creating rapport with my host family, but the English like their privacy so by the time I am ready to look at the people, they will be willing to open up as well. Hopefully.
Judging
Specifically, myersbriggs.org determines structure by asking, "In dealing with the outside world, do you prefer to get things decided or do you prefer to stay open to new information and options?" Honestly, if I stayed open to new information, I would never complete another proof. And yes, my outside world is still statistics, in case you were wondering. Of the four categories, I think this one will be the most difficult in adapting myself to England and different situations. If I am not open to new information or options, it could make the idea of a field study obsolete. Not to mention awkward if I become that horrible, America is the greatest sort of international student. However, I do not think I am like that; I just prefer to have decisions made. I am not necessarily closed to new options, ideas, or information. To be on the safe side, I will take extra care to NOT be closed minded in England. If that is indeed my default, I should be able to override it with a conscious effort to be open minded.
Source (2/24)
Woods, D.C. and S.M. Lewis. "Continuous optimal designs for generalized linear models under model uncertainty." Journal of Statistical Theory and Practice. 5 (2011): 137-145. Electronic.
With more and more complex data and experiments, linear regression is often "inadequate." Even with the "modern" regression methods such as b-splines and smoothing kernels, standard factorial designs cannot be modeled well. This papers suggests are more exact designs for specific experiments using sophisticated design selection and criterion to allow uncertainty in the link (or knot) functions. The algorithm's efficiency is tested using simulation studies.
Again, this article helps me narrow down the statistics portion of my project. Understanding previous papers written by Dr. Woods helps me create a more solid literature review, helping me to develop my topic-specific information and significance.
With more and more complex data and experiments, linear regression is often "inadequate." Even with the "modern" regression methods such as b-splines and smoothing kernels, standard factorial designs cannot be modeled well. This papers suggests are more exact designs for specific experiments using sophisticated design selection and criterion to allow uncertainty in the link (or knot) functions. The algorithm's efficiency is tested using simulation studies.
Again, this article helps me narrow down the statistics portion of my project. Understanding previous papers written by Dr. Woods helps me create a more solid literature review, helping me to develop my topic-specific information and significance.
Wednesday, February 22, 2012
Source (2/22)
Woods, Dave and Peter van de Ven. "Blocked Designs for Experiments with Correlated Non-Normal Response." Technometrics. 53.2 (2011): 173-182. Electronic.
In simple linear models, the assumptions are very strict and often unattainable. Often, "many experiments measure a response that cannot be adequately described by a linear model with normally distributed errors." The authors developed a general method of creating efficient blocked designs where the response is distributed as an exponential family using Generalized Estimating Equations. "This methodology is appropriate when the blocking factor is a nuisance variable, as often occurs in industrial experiments." Using both a systematic search and a block optimal design for a Generalized Linear Model, the results are more efficient than using an optimal GLM design. This article is useful as I am trying to clarify my statistical part of the project. This allows me to gain an idea of the type of projects Dr. Woods is involved with.
In simple linear models, the assumptions are very strict and often unattainable. Often, "many experiments measure a response that cannot be adequately described by a linear model with normally distributed errors." The authors developed a general method of creating efficient blocked designs where the response is distributed as an exponential family using Generalized Estimating Equations. "This methodology is appropriate when the blocking factor is a nuisance variable, as often occurs in industrial experiments." Using both a systematic search and a block optimal design for a Generalized Linear Model, the results are more efficient than using an optimal GLM design. This article is useful as I am trying to clarify my statistical part of the project. This allows me to gain an idea of the type of projects Dr. Woods is involved with.
Subscribe to:
Posts (Atom)
