Showing posts with label Classics. Show all posts
Showing posts with label Classics. Show all posts

19 August 2015

Archiving BMCR

Just a few days after my last post on archiving to fight link rot, Ryan Baumann (@ryanfb) wrote up his impressive efforts to make sure that all the links in the recently announced AWOL Index were archived. Since I was thinking about this sort of thing for the open-access Bryn Mawr Classical Review for which I'm on the editorial board, I figured I'd just use his scripts to make sure all the BMCR reviews were on the Internet Archive. (Thanks, Ryan!)

Getting all the URLs was fairly simple, though there was a little bit of brute-force work involved for the earlier years, before BMCR settled on a standard URL format. Actually there are still a few PDFs from scans of the old print versions which I completely missed on the first pass, but once I found out they were out there, it was easy enough to go get them. (I was looking for "html" as a way of pulling all the reviews, so the ".pdf" files got skipped.)

In the end less than 10% of the 10,000+ reviews weren't already on the Archive, but are now, assuming I got them all up there. Let me know if you find one I missed.

I'm still looking at WebCite too.

20 July 2015

Frequent Latin Vocabulary - sharing the data(base)

When I first started teaching after grad school, I did a lot of elementary-Latin instruction. I felt well prepared for this because I did my graduate work at the University of Michigan, where the Classical Studies Department has a decades-long tradition of paedagogical training. It includes people like Waldo "Wally" Sweet, Gerda Seligson, Glenn Knudsvig, and Deborah Ross. One consequence of this teaching and preparation was that I became very interested myself in Latin paedagogy and my first research output was in this direction.

In particular I started looking at the Latin vocabulary that students were learning and how that related to the vocabulary that they were reading in the texts they encountered in intermediate and upper-level classes. As I investigated this, I learned that there had been a lot of work on exactly this area not only among people studying second-language acquisition, but also in Classics circles back in the 1930s, 40s and 50s. One of the more interesting people in this area was not someone that many classicists will not know, Paul B. Diederich. Diederich had quite an interesting career, working even at that early date in what is now the trendiest of educational concerns, assessment, mainly in writing and language instruction, and eventually making his way to the Educational Testing service, ETS, which gave us the SAT.

Diederich's University of Chicago thesis was entitled "The frequency of Latin words and their endings." As the title suggests it involves determining the frequency of both particular Latin words and endings for both nouns/adjectives/pronouns and verbs. In other words, a bit of what would now qualify as Digital Humanities, Diederich of course lacked a corpus of computerized texts and he had to do this counting by hand. So he made copies of the pages of major collections of Latin works, using different colors for different genres (big genres, like poetry and prose), and then cut these sheets of paper up so that each piece contained one word. Then he counted up the words (over 200,000!) and calculated the frequencies. This biggest challenge he faced was the way his method completely destroyed the context of the individual words; once the individual words were isolated, it was impossible to know where they came from. One result of this was acknowledged by Diederich in the thesis: not all Latin words are unique. For example the word cum is both a preposition meaning "with" and a subordinating conjunction "when/after/because." This meant that Diederich needed either to combine counts for these words (which he did for cum), or label such ambiguities before cutting up the paper. As he himself admits, he did a fairly good job of the latter, but didn't quite get them all. Another decision that had to be made was what to do with periphrases, that is, constructions that consist of more than one word. Think of the many English verb forms that fall into this category: did go, will go, have gone, had gone, am going, etc. Do you want to count "did go" as one word or two?

Interesting to me was that Diederich was careful to separate words that the Romans normally wrote together. These usually short words, called enclitics, were appended in Latin to the preceding words, a bit like the "not" in "cannot" (which I realize not everyone writes as one word these days). This was a good choice on Diederich's part, as one of these words, -que meaning "and," was the most frequent word in his sample. (As a side note, some modern word-counting tools, like the very handy vocab tool in Perseus, do not do count enclitics at all. Such modern tools also can't disambiguate like Diederich could, so you'll see high counts for words like edo, meaning "eat," since it shares forms with the very common word esse, "to be." Basically we're trading off automation, and its incredible speed increases, for lack of ambiguity.)

The article I eventually published (“Frequent Vocabulary in Latin Instruction,” The Classical World 97, no. 4 (2004): 409-433) involved me using the computer to create a database of Latin vocabulary and then counting frequencies for a number of textbooks, comparing them to another set of frequent-vocabulary lists. I put some of the results of this work up on the internet (here, for example), but didn't do a lot of sharing of the database itself. This wasn't so easy way back in the early 'aughts, but it is now. Hence this post (which is a great example of burying the lede, I suppose).

I created the database in FileMake Pro version 3. Then migrated to version 6, then 8, and now 12. (Haven't made the jump to 13 yet.) Doing this work in a tool like FMP has its pros and cons—and was the subject of some debate at our LAWDI meetings a few years ago. Big on the pro side is the ease of use of FMP and the overall power of the relational-database model. On the con side is the difficulty in getting the data back out so that it can be worked on with other tools that can't really do the relational thing. For me FMP also allowed the creation of some very nice handouts for my classes, and powerful searches once I got the data into it. In the end though, if I'm going to share some of this work, it should be in a more durable and easily usable form, and put someplace where people can easily get to it and I won't have to worry too much about it. I decided on a series of flat text files for the format, and GitHub for the location. I'm going to quote the README file from the repository for a little taste of what the conversion was like:

Getting the data out of FMP and into this flat format required a few steps. First was updating the files. I already had FMP3 versions along with the FMP6 versions that I had done most of the work in. (That's .fp3 and .fp5 for the file extensions.) Sadly FMP12, which is what I'm now using, doesn't directly read the .fp5 format at all, and FMP6 is a Classic app, which OS X 9 (Mavericks) can't run directly. So hereʻs what I did:
  • Create a virtual OS X 10.6 (Snow Leopard) box on my Mavericks system. Snow Leopard was the last OS X version to be able to run the Apple Classic emulator, Rosetta. That took a little doing, since I updated the various sytem pieces as needed. Not that this version of OS X can be super secure, but I just wanted it as close as possible.
  • Convert the old .fp5 files to .fp7 with FMP 8 (I keep a few versions of FMP around).
  • Archive the old .fp5 files as a zip file.
  • Switch back to Mavericks.
  • Archive the old .fp7 files. I realized that the conversion process forced me to rename the originals, and zipping left them there, so I could skip the step of restoring the old filenames.
  • Convert the .fp7 to .fmp12.
  • Export the FMP files as text. Iʻm using UTF-16 for this, because the database uses diarheses for long vowels (äëïöü). Since this is going from relational to flat files, I had to decide which data to include in the exports.
  • Convert the diarheses to macrons (āēīōū). I did this using BBEdit.
  • Import the new stems with macrons back into FMP. I did it this way because search and replace on BBEdit is faster than in FMP.
  • Put the text files [on Github].
FMP makes the export process very easy. The harder part was deciding which information to include in which export. An advantage of the relational database is that you can keep a minimal amount of information in each file and combine it via relations with the information in other files. In this case, for example, the lists of vocabulary didn't have to contain all the vocabulary items within their files, but simply a link to those items. For exports though you'd want those items to be there for each list. Otherwise you end up doing a lot of cross-referencing. It's this kind of extra work, which admittedly can be difficult, especially when you have a complicated database that you designed a while back, that makes some avoid FMP (and other relational databases) from the start.

In the end though, I think I was successful. I created three new text files, which reflect the three files of the relational database:
  1. Vocabulary is in vocab.tab. These are like dictionary entries.
  2. Stems, smaller portions of vocal items, are in stems.tab. The vocab items list applicable stems, an example of something that was handled relationally in the original.
  3. The various sources for the vocabulary items are in readings.tab. It lists, for example, Diederich's list of 300 high-frequency items.
I also included the unique IDs that each item had in each database, so it would be possible to put them back together again, if you wanted (though you could just ask me for the files too). See the README and the files themselves for more detail. I feel pretty good though about my decision to use FMP. It was—and is, even if I'm not teaching a lot of Latin these days—a great tool to do this project in, and getting the data back out was fairly straightforward. 

You can check out the entire set of files at my GitHub repository. And here's a little article by Diederich on his career. He was really an interesting guy, a classicist working on what became some very important things in American higher education.

07 July 2015

List of undergrad Classics programs - more pandoc & github

I've kept ("maintained" would be too strong a word) a list of undergraduate classics program for a while now. I first started it when I was the webmaster for the American Classical League back in grad school and early in my career, but then I just kept it around because it seemed like a shame not too.

The other day I got a now rare update from someone on the list, and I thought this would be a good time to change the way I handle the whole thing.

First I figured I'd switch over the basic form of this very simple page from html to the easier to manage markdown, and then change my workflow to use pandoc to generate the html from this markdown whenever necessary. This part was fairly simple, though there were a few complications. For example, pandoc doesn't like html name tags, so in order to keep a bunch of internal links I had, I needed to convert those anchors to spans with IDs that matched the names I was using. BBEdit and its nice search-and-replace functionality to the rescue. I then did some minimal tweaks to the new pandoc-generated markdown file, so that pandoc could now generate some decent html from it.

Step 2 was to put the new markdown and html up on GitHub, instead of in my institutional filespace. Not only does this give me some version control over the files, but I can let other people edit the markdown and send me pull requests, instead of emailing me updates that I then put into the file. I don't think I'll get a lot of these, but you never know.

So I can now (1) make changes to the markdown myself, or accept pull requests with them in it, (2) sync my local github repo with the on-line version, (3) run pandoc on my local copy, and (4) sync it back to GitHub where I release a new tagged version. I then link to the files directly via rawgit. (The tag in GitHub is needed to make rawgit grab the correct version instead of the one it caches.)

The relevant links are here:


Check out the list and let me know about any updates via a pull request!

03 November 2013

Teaching with ORBIS #lawdi

On the last day of our Linked Ancient-World Data Institute this summer, sponsored by the NEH Office of Digital Humanities, I argued that it was important to show how all this exciting work could have practical implications for what professional classicists (and other ancient-world types) spend a lot of time on, teaching. To that end, I promised to do a post describing how I had assigned my Classical-archaeology students some short homework using Stanford's great new tool, ORBIS. What's ORBIS? In the words of the site, ORBIS "reconstructs the time cost and financial expense associated with a wide range of different types of travel in antiquity." More simply it allows you to map routes between two places in the Roman world given certain constraints for cost, time, and type of route.

Naturally the fine people at ORBIS did a nice upgrade to the service after I made that promise, but before I got it completed, so I had some more work to do before this post. (Fair enough, I dragged my feet for too long anyway. Nemesis strikes!) But finally here it is, suitable for framing (or at least bookmarking).

A Very Short Guide to Using ORBIS in your Classical-Archaeology Course

1. RTFM

Make sure that you understand as much as possible about ORBIS, what it is, how it works, and so on. You don't need to be an expert, but you should at least be able to do more than your students will by the time they finish the assignment. It won't take more than an hour to read the "Introduction to ORBIS", "Understanding ORBIS" and "Using ORBIS" tabs on the website. Don't miss the nifty how-to videos. Although there's more there to read, these three sections will get you far enough for step 2.

2. Make sure you know how to use it

Play around a bit yourself on the "Mapping ORBIS" section. Try to get from one place to another. Change the various parameters. Use all the controls, so you know how to change the views of the route and so on. Click the buttons and links and sliders. Go nuts. Depending on your technological prowess, this will take you a few hours at most.

3. Demo it in class

Once you're confident that you can show your students the basics of the site with confidence, have your student read those same three sections of the ORBIS website that you read up in #1 in preparation for a short demo of ORBIS that you'll do in class for them. Nothing fancy, just enough to show them the basics. I like to point out to them how long travel takes when you don't have motorized vehicles and how much faster travel over sea is than over land, but be sure to walk them through creating a route and choosing the various options, no matter what extra details you cover.

4. Assignment 1 of 2

Have your students use ORBIS to find a simple route between two places that you specify. Have them do it under multiple conditions. (I used three different sets.) Then have them either print out or take a screen shot of the result, with all routes shown. Here's one with routes between three sets of cities, one taking the fastest route, another the cheapest, and the third the shortest.
Since you've set the parameters, you'll know what the correct routes should be, and thanks to the different colors, it's easy to tell at a glance whether the student got it right. Successful completion of this will indicate that your students can handle using the basics of ORBIS. Make sure they all successfully complete this first assignment. Then they're ready for part 2.

5. Assignment 2 of 2

This part is up to you. Depending on which section of your course you're using ORBIS in, you'll want to find some question you can answer, or some issue you can illuminate via ORBIS. I actually did something that was completely out of ORBIS' chronological span.

To help my students understand the rationale behind some of the placement of early Greek colonies in the west, I had them examine routes between Delphi and Naples. The latter was used as a proxy for the earliest colony of Pithecoussae. (This was actually a variation on an assignment I had made up years ago using a QuickTime movie with links to the Perseus website.) The biggest travel difference between the later Roman empire, the time in which ORBIS is "located", and the geometric period was the roads in use. Obviously none of the vast Roman road system was in place, and so I made sure the students used only routes that avoided long portions over land. (I want to use this assignment again, so I'm not giving away all the details!)

6. Put the students to work

If you subsequently have your students come up with their own ORBIS projects, odds are they'll find something useful and interesting to do, perhaps something you hadn't quite thought of. A set of mine, for example, used ORBIS to explore the different travel experiences of three characters from the ancient world with differing socio-economic backgrounds, complete with clever backstories!


And there it is. Hope this encourages you to use ORBIS and other terrific ancient-world-related DH tools in the classroom! I'd love to hear in the notes about your experiences with ORBIS or with any other tool.

04 October 2012

Update II: "'Crisis' in Classics" briefly revisited

The APA has just announced the results of the annual elections for 2012. So what kind of institutional representation do we find the members voted for (not that there was much choice in this dimension)?

Office              Institution     Description
President           UCincinnati     Big Public U.
Financial Tr.       UC, Davis       Big Public U.
VP Prof. Matters    UVa             Big Public U.
VP Pub & Research   UT              Big Public U.

Board of Directors  Ohio State      Big Public U.

Board of Directors  UPenn           Big Private U.
Nominating Comm     Princeton       Big Private U.
Nominating Comm     CUNY            Big Public U.
Education Comm      Phillips-Exeter Elite Prep School

Goodwin Award Comm  Bowdoin         Elite Private LAC
Prof Matters Comm   Episcopal Acad  Elite Prep School
Program Committee   Harvard         Big Private U.

Pub & Research Comm Cornell         Big Private U.

So that's 6 from big public universities, 5 from big private universities, 2 from elite prep schools (1 of which has an endowment that easily dwarfs that of all but the wealthiest of LACs in this country) and 1 from an elite small LAC.

Classics' main professional organization continues to be dominated by the "haves."

Reminds me of the Romney campaign...


Previous posts in this category:

My New article

I've been getting more and more involved in digital humanities, so I'm happy to announce that my most recent publication is now out in an open-access journal:
John D Muccigrosso, “Re‐Interpreting the Robinson Skyphos,” Studia Humaniora Tartuensia 13, no. A.1 (2012): 1–15
That's the old-school citation for this new-school journal, and hardly befitting a 21st-century open-access publication, so here are a couple of better choices:
  1. The direct link to the journal webpage for the article
  2. My Zotero library reference
  3. The article itself as a pdf
The article is of course free to download (hence the OA bit), so knock yourself out. Please.

Here's the abstract:
The scene on the Robinson skyphos was wrongly identified for years as a depiction of clay‐working, either in a kiln or other preparation area. Recent scholarship has correctly identified it instead as one related to the grain harvest. This article presents a new examination of the scene, pointing out details the importance of which had not previously been noted. It also brings to bear comparanda from Egyptian art which put the identification of the scene beyond doubt.
The article began as part of an exploration of depictions of what were called potters and pottery workshops on ancient Greek pots, but which I thought were often not. It's inspired to a large extent by the work of David Gill and Michael Vickers on the elevation of ancient pottery-making to an "art" instead of a "craft" to reflect modern rather than ancient thinking. (And let's not get into the whole issue of how we distinguish between those two things!) The work actually started as an grad-school exploration of how much Greek pottery was exported, not only in terms of the number of physical pots, but also the economic value of those pots. It won't be surprising to learn that the course was taught by William Loomis, the guy who studied how much the ancient Greeks actually got paid (Wages, welfare costs, and inflation in classical Athens), and, upon reflection, the topic was probably a good indicator of my interest in what we now call digital humanities in the first place!

And for a little academic genealogy...David Moore Robinson, the classical art historian and collector after whom the pot in the article was named, was the teacher of George Hanfmann, who in turn taught John G. Pedley, with whom I studied at Michigan (though he was not my advisor) and who invited me to join the on-going excavations at Paestum, Italy at the end of the last century(!), which were conducted by Jim Higginbotham of Bowdoin College, with whom my previous article was co-authored. So far, so good. Fairly normal academic stuff, especially for a fairly small field like mine. But wait, there's more!

One of the standard works on the manufacturing techniques of ancient Greek pottery was written by Joseph V. Noble, who died in 2007, during the period in which I was working on this article. Turns out he had lived for years in the same town as me (Maplewood, NJ), just a few hundred meters from the train station where I daily stood for my commute, though unfortunately I never knew that until it was too late.

01 February 2012

AIA Comes out in Favor of the Research Works Act

In the middle of the holiday break, our own AIA, the Archaeological Institute of America, submitted to the  Office of Science and Technology Policy of the US government their statement on the recently proposed Research Works Act, which is in essence an attack on the growing Open Access movement. (Follow the link to Thomas.gov and check out the Wikipedia entry too.)

Leaving aside the apparent absence of this document from the AIA's own website (site search engines can be remarkably crappy when it comes to this kind of thing), why didn't they think this would be worth letting me, a member in good standing, know about? Especially now that the AAA's response—to which the AIA explicitly refers in their document—has raised a ruckus in that group!

But more importantly, where's the membership on this? Are we in favor of this stance? I'm certainly not. Anyone else?

Many of us in the profession are advocates for Open Access (a term which the AIA doesn't even seem to understand, to judge from their response), and I suspect would have a thing or two to say about the stance of our professional organization. Others have made the case already, so I won't re-argue it here, but I encourage you to read some of them (by, e.g., Kristina Kilgrove or Derek Lowe).

What's most galling though, is that this statement was made by the AIA literally days before our  annual meeting, when it would have been a trivial matter to bring up the subject in official venues and get some important feedback. I wasn't there (off in Rome with students), but I haven't had any official word of anything. And given the decades-long prominence of some of our members in what's now known as the Digital Humanities, this is profoundly disappointing.

I certainly hope that others in the AIA feel the same way about this, and I'm fixing to find out who they are!

Correction: After some discussion with a few others, including Sebastian Heath, I have to correct myself. This AIA's letter was not in response the RWA per se, but, as I wrote, to the RFI from the OSTP. The issues are the same, in that the RWA addresses the question of mandating Open Access to publications dealing with federally funded research which is what the AIA statement dealt with (along with some other things I disagree with). I'll deal with this more in another post, but I wanted to get a correction in right away and apologize for the error.

11 December 2011

Dude, where's my diss? ''

Part III, in which I produce the document

In the second installment of this multi-post topic, I ended wondering where I might store copies of my dissertation for public download. For some reason it hadn't occurred to me to use the Box account that I have. (Box is like DropBox.) Since I have now figured this out, I present below two pdf versions. This first is to a copy of the UMI version which is, as I wrote before, essentially a photocopy of the original paper version I submitted to them back in 1998. The second pdf is a searchable version I recently created. That wasn't as easy as it sounds.

It's true that it's a trivial matter to create a pdf these days. On my Mac, I can just print directly to pdf, and this was my first approach. The problem is that the latest version of Microsoft Word renders the text slightly differently from the way version 5.1a (of blessed memory) did it. As a result the page numbering got way off. (I will refrain from the obvious rant about the problems this version issue causes.) The first remedy I tried was tweaking the margins a bit, thinking that the fonts (mainly Times) were being rendered at a consistently different width. No dice. In some cases lines were longer, in other shorter. I haven't a clue why. OK, I think, so I'll just fire up version 5.1a. Well, that requires at least Classic, which doesn't run on Intel Macs anymore. No problem, SheepSaver emulates such a machine, even on my nifty new MacBook Pro. First new problem: OS 9 doesn't allow such easy printing to pdf. Solved with PrintToPDF, which creates a virtual Chooser (remember that?) printer that really sends output to a pdf file. Great. The second problem wasn't new nor was it so easily solved.

Word 5.1a does a better job than the 2011 version at reproducing the layout of my original document, but not a perfect one. For some reason it was just not matching up and once again it wasn't a simple matter of adjusting margins. So here's what I did. I figured that the smallest unit of text I had to worry about was the page, and many of them were the same, that is, they started and ended on the same word as my original dissertation printout. Some of the intervening lines look different, but since no one was going to be citing my dissertation that way, it would be OK. Where the pages didn't line up, I went in and inserted extra spaces to force line breaks, with the occasional tweak to margins, mainly in indented quotations. That got the pages right, and let my virtual 1998 Mac create a searchable pdf.

The only remaining problem with the pdf is that the text in ancient Greek is not real text. Back in the 90s we still weren't using Unicode everywhere, so the ancient Greek is really just regular Latin character codes shown in a font that uses Greek glyphs instead of Latin ones. (In reality lots of the accented Greek characters are punctuation of some kind.) The pdf displays the font fine, but it really isn't Greek text that you can copy or search for.

Here are the links. Again, they lead to my Box account, which I haven't upgraded to allow direct downloads, so you'll have to do something else to get the pdf itself:


What I'd really like to do is make the dissertation available as an e-book of some kind. The problem remains the Greek and the page breaks. The Greek isn't a big deal, even if I had to type it all out agin (which I don't); there's not a lot of it. Also it's not difficult to turn a pdf into one of the popular e-book formats, but my footnotes mean I can't do that without some work. Ideally I'd start from the Word doc, so a little research is needed to see what the options are.

Meanwhile...where's your diss?

01 November 2011

"'Crisis' in Classics" briefly revisited

Back before Who killed Homer?, John Heath wrote a somewhat contentious article for Classical World (John Heath et al., “Self-Promotion and the ‘Crisis’ In Classics [with responses],” CW 89, no. 1 (1995): 3-52) in which he argued, among other things, that the significant great divide in the profession was between those elite who work at large schools and the others who work at small. In support of the argument Heath noted differing participation rates in professional organizations, which he attributed to the ability of people at larger, more resource-rich places to volunteer for such work, knowing that their institutions would provide the needed back-up. As someone working at a small, resource-poor university, I was sympathetic to that argument, though several of the respondents to his article were not.

Being a data-based kind of guy, it seems to me that this is a testable hypothesis. If the participation rates in professional organizations present the kind of skew towards larger institutions (universities) that Heath describes, that's evidence in favor of his proposition. A recent announcement of the latest officers of the major professional organization for Classicists, the American Philological Association (one of many APAs) provides one relevant data set:


President-Elect: Denis Feeney (Princeton U)
Vice President, Outreach: Mary-Kay Gamel (UCSC)
Vice President, Publications: Michael Gagarin (UTAustin)
Board of Directors: Sara Forsdyke (UMich) and Matthew Roller (Johns Hopkins U)
Nominating Committee: Donald J. Mastronarde (UCBerkeley) and Ruth Scodel (UMich)
Education Committee Member: Mary C. English (Montclair State U)
Goodwin Award Committee: Peter T. Struck (UPenn)
Professional Matters Comm. Members: Lillian Doherty (UMD) and Barbara K. Gold (Hamilton College)
Program Committee: Member Christopher A. Faraone (UChicago)
Publications Committee Member: Andrew M. Riggsby (UTAustin)

By my count, that's 12 of 13 from universities, including two from each of two places (one of which, for the record, is my doctoral alma mater in Ann Arbor), and 9 of the 10 universities have graduate programs in Classics.

I did a count a few years ago of Classics faculty by the highest degree offered at their institution. It was based on the 2002-2003 APA departmental survey data. Here's a graphical representation. You'll see that faculty at BA-offering institutions make up half of the total, with those at institutions without even a Classics major accounting for another 8%. That total of 58% is a far cry from the 15.4% (2/13) rate seen in the election results.

Lest one think this an unusual year, holders of these same offices in 2011 were employed by BU, UMD, Georgetown U, UChicago, UCLA, Stanford U, U South Carolina, Reed College, Columbia U, Duke, UPenn,* UPenn, and Brown U, respectively. That's one college out of 13 positions, two from one university,* and again two from places without graduate programs.

Of course not all universities are the same (I teach at one, for example, though it's very much a small liberal-arts college attached to a theological school), and different faculty have different workloads, but the universities in the lists above are large and in most cases wealthy institutions. Add to that the information about graduate programs, I'd say this makes a reasonable prima facie case in support of Heath's position.

* I had to cheat a bit with one of the committee memberships because they were two members elected for 2012, but only one for 2011, so I went back a grabbed the previously elected member.