This thing gets indexed by Google, so I'll post this here for the next person who wonders. (I've also updated some random wiki page I wound when googling, and should probably mail the soekris-tech list while I'm at it, but probably won't do that tonight.)
Anyway. The Soekris net4526 single-board computer has on it an 8-bin header connector labelled JP3, which is known to carry five GPIO lines. However, Soekris Engineering hasn't released a manual for the net4526, only for the related net4501, and the only related traffic I found on their mailing list was someone who had figured out some of the pins and was asking for help with the others. (Which is odd, because I thought I'd seen a complete pinout posted there at one point.)
So, since I have a project that might need this feature, I went in and tested myself, and (arranged like they are here, with the JP3 label upright):
1: +3.3V 2: +5V
3: GPIO 7 4: GPIO 8
5: GPIO 11 6: GPIO 21
7: GPIO 22 8: GND
I saw some talk about GPIO numbers in octal or something, so I will add: those are in decimal and zero-indexed, as accepted and reported by NetBSD's gpioctl(8) command. ALso, I have independently confirmed those of the pin assignments noted in the mailing list message I mentioned.
Know further that the pins' input mode has an internal pullup.
The connector I ginned up for the header is also… interesting, and may deserve a post of its own.
2009-09-04
2009-07-14
The truth is also a three-edged sword.
But these are a different three from the canonical ones for understanding. Of course I'm talking about the RAIDframe stuff, because that's all I seem to get around to posting here. Anyway, there's what we believed at boot, what we know now, and what we want to be believed on the next boot. For the code I've written, they're fields of the same struct. For the existing raid(4) code, the information can be a bit more… scattered. (Making things slightly more fun: all the metadata for a RAID set is replicated on each component, so there's the question of what to do if there are non-fatal differences.) My SoC mentor has noted that things could use some reorganizing there, and part of me would like that too, but a much larger part of me says It's Working Code, Leave It Alone.
This notion applies to fault-tolerant persistent data systems in general, really; there, as with RAIDframe, the first item is relevant only until some kind of roll-forward is done to clean up after a failure. In RAIDframe's case, this is raidctl -P, and it's a little more prominent because it runs in parallel with, to borrow a term, the mutator; contrast with wapbl(4)'s roll-forward, which is done automatically on mount, appears to be quite quick, and I assume blocks use of the filesystem until it's done.
(In the other half of my life I'm looking at the literature on persistence, and it's almost odd how these two things are converging here.)
This notion applies to fault-tolerant persistent data systems in general, really; there, as with RAIDframe, the first item is relevant only until some kind of roll-forward is done to clean up after a failure. In RAIDframe's case, this is raidctl -P, and it's a little more prominent because it runs in parallel with, to borrow a term, the mutator; contrast with wapbl(4)'s roll-forward, which is done automatically on mount, appears to be quite quick, and I assume blocks use of the filesystem until it's done.
(In the other half of my life I'm looking at the literature on persistence, and it's almost odd how these two things are converging here.)
2009-06-12
It Is Written
What I'll call a first draft of the RAIDframe parity map stuff is written, and compiles and links, and if run will actually do something. That something will, in practice, probably involve bugs.
Now to get my QEMU setup into something more resembling a useful state, because the time I spend on that will almost certainly be paid back by not waiting for my test box to reboot, and I've been meaning to deal with that anyway. Once this gets to the point of serious benchmarking I'll need to use actual hardware for the most part, of course.
The RAIDframe codebase, incidentally, is… not unelaborate.
Now to get my QEMU setup into something more resembling a useful state, because the time I spend on that will almost certainly be paid back by not waiting for my test box to reboot, and I've been meaning to deal with that anyway. Once this gets to the point of serious benchmarking I'll need to use actual hardware for the most part, of course.
The RAIDframe codebase, incidentally, is… not unelaborate.
2009-06-11
The RAID Project: Things Not To Do
- Let writes to the RAID hit the disk before the corresponding parity map bit is set on disk.
- Let writes to the RAID hit the disk after the corresponding parity map bit is cleared on disk. (That is, updates which just mark regions clean again still need a barrier.)
- Have one write see that its region needs to be marked unclean, then do that, and then before that actually gets committed to the disk another write to the same region sees that it's allegedly already marked and just does the write, which happens to hit the disk before the parity map update and then the power goes out at that exact moment.
This may not even be possible — I think I'll only ever be starting writes from one particular thread, given the RAIDframe architecture, though I'm not sure of that yet — and even if it is it sounds stunningly unlikely. Which is to say that if I get this wrong I may never find out; so don't do that.
Point also being that it's important to keep invariants in mind when dealing with shared-state concurrency, including those invariants that involve the state of secondary storage and the potential behavior of loosely specified hardware as well as the program's data structures proper.
Completely unrelatedly, I've just learned that the posting interface here rejects ill-formed HTML.
2009-06-01
Parity
I didn't go into too much detail about my Google Summer of Code project last time. It is: improving RAIDframe parity handling. And now I'm going to be excessively verbose about it.
Specifically: the thing about RAID levels that provide redundancy (i.e., not RAID 0) is that there's some kind of invariant over what's on the disk: both halves of a mirror are the same, or each parity block is the XOR of its corresponding data blocks, &c. And the thing about software RAID is that, if the power goes out (or the system crashes) while you're in the middle of writing stuff to each of the disks, some of those writes might happen while others don't. Then, when the lights come back on, the invariant may no longer hold for any stripe that was being written.
This is of particular concern for RAID 5, because if the parity is still wrong when (not if) a disk fails and one of the data blocks needs to be reconstructed by XORing the parity with the remaining data, you will get complete garbage instead of the data you lost. This is bad.
One solution, and the one currently used in NetBSD, is to set a flag on each disk making up the RAID when it's configured, and clear it when it's unconfigured. If that flag is already set when the set is brought up, then there might have been an unclean shutdown requiring the parity to be recomputed.
That is, requiring the entire array to be read from beginning to end. Which, as magnetic disk drives pack more and more tracks onto their platters, inevitably takes longer and longer. As it is, each unclean shutdown requires many hours of parity rewriting, during which the disk I/O load interferes with whatever the system's actual job is. This is also kind of bad.
It is said that the Solaris Volume Manager (which I briefly administered an instance of, but didn't have to care how it worked in this much detail) divides the RAID into some number of regions and records for each one whether its parity might be out of sync. This seems like a simple enough idea.
Except it's kind of not. Ideally, you'd like as many of these regions to be marked clean as possible, to cut down on the parity rewriting time. On the other hand, because you'll have to do disk seeks (and probably disk cache flushes, too, and hope the firmware isn't too broken) to set or clear a region's dirty bit, and it's absolutely essential that that bit-setting hit the disk before any writes to the region are done, you also want to hold off on marking clean those regions that you think might be getting written to sometime soon.
So, if you're getting truly random I/O, then you're kind of stuck. But, if what's on top of the RAID is some halfway reasonable filesystem that's been painstakingly designed to exhibit reasonable locality of reference, then recent write activity should (at the region level) be a decent predictor of the future. I hope.
And then there's the part of the project where I get all this integrated into the kernel, which is beyond the scope of this post.
Specifically: the thing about RAID levels that provide redundancy (i.e., not RAID 0) is that there's some kind of invariant over what's on the disk: both halves of a mirror are the same, or each parity block is the XOR of its corresponding data blocks, &c. And the thing about software RAID is that, if the power goes out (or the system crashes) while you're in the middle of writing stuff to each of the disks, some of those writes might happen while others don't. Then, when the lights come back on, the invariant may no longer hold for any stripe that was being written.
This is of particular concern for RAID 5, because if the parity is still wrong when (not if) a disk fails and one of the data blocks needs to be reconstructed by XORing the parity with the remaining data, you will get complete garbage instead of the data you lost. This is bad.
One solution, and the one currently used in NetBSD, is to set a flag on each disk making up the RAID when it's configured, and clear it when it's unconfigured. If that flag is already set when the set is brought up, then there might have been an unclean shutdown requiring the parity to be recomputed.
That is, requiring the entire array to be read from beginning to end. Which, as magnetic disk drives pack more and more tracks onto their platters, inevitably takes longer and longer. As it is, each unclean shutdown requires many hours of parity rewriting, during which the disk I/O load interferes with whatever the system's actual job is. This is also kind of bad.
It is said that the Solaris Volume Manager (which I briefly administered an instance of, but didn't have to care how it worked in this much detail) divides the RAID into some number of regions and records for each one whether its parity might be out of sync. This seems like a simple enough idea.
Except it's kind of not. Ideally, you'd like as many of these regions to be marked clean as possible, to cut down on the parity rewriting time. On the other hand, because you'll have to do disk seeks (and probably disk cache flushes, too, and hope the firmware isn't too broken) to set or clear a region's dirty bit, and it's absolutely essential that that bit-setting hit the disk before any writes to the region are done, you also want to hold off on marking clean those regions that you think might be getting written to sometime soon.
So, if you're getting truly random I/O, then you're kind of stuck. But, if what's on top of the RAID is some halfway reasonable filesystem that's been painstakingly designed to exhibit reasonable locality of reference, then recent write activity should (at the region level) be a decent predictor of the future. I hope.
And then there's the part of the project where I get all this integrated into the kernel, which is beyond the scope of this post.
2009-05-21
The blog is alive!
This summer, I'm participating in Google's Summer of Code, attached to the NetBSD project and working on fixing certain misfeatures of the software RAID driver — more details on which later.
Various other SoCers have declared their intention of keeping a weblog on their work; and, since I'm not entirely unfamiliar with the concept, I thought I might do that as well. And this blog, being created for me to do Serious Technical Blogging (and then left to gather dust when I stopped being bothered with writing for it), and hosted by Google no less, seems like the best place.
Well. I briefly considered “blogging” by hand-writing an RSS file and giving out that URL — why, some web browsers even format RSS nicely when you browse to it — but then came to my senses. I also considered making a separate blog (as BlogSpot has a nice interface/ontology for that), but didn't really see the point. And, hey, tags; they're what Web 2.0 is all about, except when it's not, or something.
Various other SoCers have declared their intention of keeping a weblog on their work; and, since I'm not entirely unfamiliar with the concept, I thought I might do that as well. And this blog, being created for me to do Serious Technical Blogging (and then left to gather dust when I stopped being bothered with writing for it), and hosted by Google no less, seems like the best place.
Well. I briefly considered “blogging” by hand-writing an RSS file and giving out that URL — why, some web browsers even format RSS nicely when you browse to it — but then came to my senses. I also considered making a separate blog (as BlogSpot has a nice interface/ontology for that), but didn't really see the point. And, hey, tags; they're what Web 2.0 is all about, except when it's not, or something.
2008-06-19
Vocabulary Time!
ajax, v.t.: Of a dynamic website, to cause a user to take an unintended action (esp. one difficult or impossible to undo) because the interface rearranged itself (e.g., by removing an item from a list) in delayed response to an earlier action, concurrently with the user's attempting to activate an interface element. Often used passively, as e.g. I just got ajaxed!
2008-06-04
Before It Was Popular?
I remember when I heard about DieHard, the probabilistic memory-safety contrivance, at a talk at NEPLS; that was in 2005, so their PLDI paper wouldn't have existed yet.
And then today I see them get name-checked on hack-a-day — perhaps because they've taken to promoting it as a way to increase the security of the popular web browser Firefox.
Go figure.
And then today I see them get name-checked on hack-a-day — perhaps because they've taken to promoting it as a way to increase the security of the popular web browser Firefox.
Go figure.
2008-04-21
Damn you, Microsoft!
Many people have reason to curse Microsoft, of course, but this is a bit different.
Microsoft's greatest and perhaps only contribution to humanity, the Trackball Explorer, seems to have become a collector's item (with second-hand prices to match!) while I wasn't looking.
Mine is still working (for now), but I'd wanted to have one at home and one at school (read “in the office”), like I did at my last job, which is where I became acquainted with this device, and now that can't happen.
If only I'd known or even suspected back in 2005 or so; I would have picked up a few more.
But perhaps it's only appropriate, given that my extinct trackball is paired with an IBM Model M keyboard which marked its 21st birthday last Thursday. (It was not celebrated, as I'd misremembered the date. Oops.) This means it can legally drink alcoholic beverages now.
Microsoft's greatest and perhaps only contribution to humanity, the Trackball Explorer, seems to have become a collector's item (with second-hand prices to match!) while I wasn't looking.
Mine is still working (for now), but I'd wanted to have one at home and one at school (read “in the office”), like I did at my last job, which is where I became acquainted with this device, and now that can't happen.
If only I'd known or even suspected back in 2005 or so; I would have picked up a few more.
But perhaps it's only appropriate, given that my extinct trackball is paired with an IBM Model M keyboard which marked its 21st birthday last Thursday. (It was not celebrated, as I'd misremembered the date. Oops.) This means it can legally drink alcoholic beverages now.
2008-04-10
Negativity
On Tuesday, I stayed home sick with a cold-and/or-flu (thus missing a class that might have held some information I hadn't yet gotten from the original paper); but people on the Internet were arguing about religion. As they often do. And, as they often do, someone said “it’s true technically you can’t prove a negative”.
This is, of course, false. Negatives can be proven, even without the law of the excluded middle, and even without your logic having a primitive concept of “not”. In, say, the Calculus of Inductive Constructions, “not P” is “P → False”, where “False” is an inductive predicate with no constructors, so you get your ex falso quodlibet just from case analysis, and it's all terribly pretty.
Anyway. Negatives. Even without this newfangled lambda-cube stuff, you can prove that the square root of 2 is irrational, which is to say that there do not exist integers p and q, where q ≠ 0, such that p2 = 2 q2. This was first done around 500 BC, and the person responsible was drowned. Modern Internet religion arguments are, thankfully(?), much tamer.
So, last year, when I was first learning about the Coq proof assistant, I used it to obtain a machine-checkable proof of just that negative. Knowledgeable readers will notice that I independently reinvented part of the NArith library there, because I didn't know about the original.
Thus, what better to do on such a Tuesday, while feverish and generally out of it, but fiddle with an esoteric logic? Which is why I now have a much shorter proof, though I didn't bother mapping it back to the Peano numbers like I did with the other one.
The next obvious thing is, of course, to prove irrational the square roots of other numbers than 2. Or all of them! Or at least all the primes, because it's only in prime bases that the count of trailing zeroes (ctz ≈ bsf) interacts with multiplication in the way needed. (Hints: 2*5=10, 9*9=81.) And that I eventually gave up on.
But, first, I defined subtraction using specification types: given natural numbers a and b, return either c and a proof that a = b + c, or a proof that a < b. And then division: given n and d and a proof of d > 0, return either q and a proof that n = d * q, or a proof that no such q exists.
That last one sounds trivial, but of course it isn't; there's no excluded middle, but, perhaps more to the point, the proof of that “it is or it isn't” can be converted into an actual, runnable, proven-correct OCaml program which does the division. Not that that's very practically useful — it works on Peano numbers and does repeated subtraction (which in turn is repeated decrement).
But it does mark the first time I've successfully used Coq's support for strong induction. Twice, for the division and then for the count-trailing-zeroes, which... or, maybe instead of trying to gloss it into English, I'll just post the original type:
This is, of course, false. Negatives can be proven, even without the law of the excluded middle, and even without your logic having a primitive concept of “not”. In, say, the Calculus of Inductive Constructions, “not P” is “P → False”, where “False” is an inductive predicate with no constructors, so you get your ex falso quodlibet just from case analysis, and it's all terribly pretty.
Anyway. Negatives. Even without this newfangled lambda-cube stuff, you can prove that the square root of 2 is irrational, which is to say that there do not exist integers p and q, where q ≠ 0, such that p2 = 2 q2. This was first done around 500 BC, and the person responsible was drowned. Modern Internet religion arguments are, thankfully(?), much tamer.
So, last year, when I was first learning about the Coq proof assistant, I used it to obtain a machine-checkable proof of just that negative. Knowledgeable readers will notice that I independently reinvented part of the NArith library there, because I didn't know about the original.
Thus, what better to do on such a Tuesday, while feverish and generally out of it, but fiddle with an esoteric logic? Which is why I now have a much shorter proof, though I didn't bother mapping it back to the Peano numbers like I did with the other one.
The next obvious thing is, of course, to prove irrational the square roots of other numbers than 2. Or all of them! Or at least all the primes, because it's only in prime bases that the count of trailing zeroes (ctz ≈ bsf) interacts with multiplication in the way needed. (Hints: 2*5=10, 9*9=81.) And that I eventually gave up on.
But, first, I defined subtraction using specification types: given natural numbers a and b, return either c and a proof that a = b + c, or a proof that a < b. And then division: given n and d and a proof of d > 0, return either q and a proof that n = d * q, or a proof that no such q exists.
That last one sounds trivial, but of course it isn't; there's no excluded middle, but, perhaps more to the point, the proof of that “it is or it isn't” can be converted into an actual, runnable, proven-correct OCaml program which does the division. Not that that's very practically useful — it works on Peano numbers and does repeated subtraction (which in turn is repeated decrement).
But it does mark the first time I've successfully used Coq's support for strong induction. Twice, for the division and then for the count-trailing-zeroes, which... or, maybe instead of trying to gloss it into English, I'll just post the original type:
Definition ctz_spec : forall n b, 1 < b -> { e |
(exists2 m, 0 < m & n = expt b e * m) &
(forall e' m, e < e' -> 0 < m -> n <> expt b e' * m) } + { n = 0 }.
The Spammers Have Won
My last post here warned, in a cryptic and laconic way, of the dangers of trusting automated processes over human judgment — to wit, it was a Google search for the result of people doing a search-and-replace on the font name “Arial” to replace it with the font name “Verdana”, but also altering content as well as markup, and partial words as well as whole. Thus, “adversarial” → “adversverdana”. Cute, huh? (No, I don't go trolling for that kind of thing. I was reading one of the documents so afflicted.)
Apparently, between that and the way that I'll leave this thing unattended for months while meaning to post stuff but not actually doing that thing (in part because, well, originally I'd been hoping to write actual halfway-decent posts here, not twitters-out-of-place like that), I've been declared a possible spam blog and must now transcribe letters from an image to prove my humanity in order to make or edit posts.
Really? In my last job, I got paged in the middle of the night, on weekends, on weekends in the middle of the night, and even on vacation HOW many times, because some mail system had fallen over with its legs in the air from too much spam? And I spent HOW much time delicately tuning Postfix's rate-limiting, because there just wasn't enough hardware for the spam, but I didn't want to delay actual people's LKML subscriptions or whatever as collateral damage? To get this.
So, on the one hand, I can understand where the Google Overlords are coming from, but on the other hand, I'm kind of insulted. And I usually don't take computers' opinions of me all that personally.
Meanwhile, I note that, mysterious spam flag or no, infrequently updated or no, this blog still comes up in the first page of a Google search for... my name. Which, as it's shared by someone famous enough to have his own Wikipedia page, is not a completely vacuous achievement. Or something.
Apparently, between that and the way that I'll leave this thing unattended for months while meaning to post stuff but not actually doing that thing (in part because, well, originally I'd been hoping to write actual halfway-decent posts here, not twitters-out-of-place like that), I've been declared a possible spam blog and must now transcribe letters from an image to prove my humanity in order to make or edit posts.
Really? In my last job, I got paged in the middle of the night, on weekends, on weekends in the middle of the night, and even on vacation HOW many times, because some mail system had fallen over with its legs in the air from too much spam? And I spent HOW much time delicately tuning Postfix's rate-limiting, because there just wasn't enough hardware for the spam, but I didn't want to delay actual people's LKML subscriptions or whatever as collateral damage? To get this.
So, on the one hand, I can understand where the Google Overlords are coming from, but on the other hand, I'm kind of insulted. And I usually don't take computers' opinions of me all that personally.
Meanwhile, I note that, mysterious spam flag or no, infrequently updated or no, this blog still comes up in the first page of a Google search for... my name. Which, as it's shared by someone famous enough to have his own Wikipedia page, is not a completely vacuous achievement. Or something.
2008-04-01
2008-02-28
Fun with the Value Restriction
This is not a valid OCaml program:
The typechecker rejects it, with this error message:
Removing the last line makes it typecheck. As does exchanging the second and third lines, thusly:
Yes. Notice the mysterious action-at-a-distance. Notice also how the type inferred for poly (see ocamlc -i) is not polymorphic at all, but rather thing endo, the same as mono. Notice further that it would not be possible to define it with an explicit ascription of that type before the definition (i.e., outside the scope) of thing itself. Apparently this is bad.
And then there's the reason why the type of poly didn't have its type variable generalized, thus leaving it with the type '_a endo and subject to unification with the type of mono. This is, of course, the value restriction of ML fame, which I will not explain here; but it's a little more confusing (read: I spent far too much time trying to come up with the mostly minimal test case seen here) (so, yes, this actually is something I ran into in a halfway-real program) in OCaml, where it's been relaxed in certain ways. But not others. For example, this:
Is fine.
type 'a endo = Endo of int * ('a -> 'a)
let poly = Endo ((succ 23),(fun x -> x))
type thing = { u : unit }
let mono = Endo (0,(fun { u = u } -> { u = u }))
let _ = [mono;poly]The typechecker rejects it, with this error message:
File "vrmh.ml", line 5, characters 14-18:
This expression has type 'a endo but is here used with type thing endo
The type constructor thing would escape its scopeRemoving the last line makes it typecheck. As does exchanging the second and third lines, thusly:
type 'a endo = Endo of int * ('a -> 'a)
type thing = { u : unit }
let poly = Endo ((succ 23),(fun x -> x))
let mono = Endo (0,(fun { u = u } -> { u = u }))
let _ = [mono;poly]Yes. Notice the mysterious action-at-a-distance. Notice also how the type inferred for poly (see ocamlc -i) is not polymorphic at all, but rather thing endo, the same as mono. Notice further that it would not be possible to define it with an explicit ascription of that type before the definition (i.e., outside the scope) of thing itself. Apparently this is bad.
And then there's the reason why the type of poly didn't have its type variable generalized, thus leaving it with the type '_a endo and subject to unification with the type of mono. This is, of course, the value restriction of ML fame, which I will not explain here; but it's a little more confusing (read: I spent far too much time trying to come up with the mostly minimal test case seen here) (so, yes, this actually is something I ran into in a halfway-real program) in OCaml, where it's been relaxed in certain ways. But not others. For example, this:
type 'a endo = Endo of int * ('a -> 'a)
let poly = Endo (24,(fun x -> x))
type thing = { u : unit }
let mono = Endo (0,(fun { u = u } -> { u = u }))
let _ = [mono;poly]Is fine.
2008-01-27
Wheels Within Wheels
Exhibit A: Hashed and Hierarchical Timing Wheels: Efficient Data Structures for Implementing a Timer Facility, by George Varghese and Tony Lauck, originally in SOSP '87 and later reappeared in 1996 when network protocol research had caught up to it.
Exhibit B: An implementation of hierarchical timing wheels for the fleshy-ape platform, due to David Allen (2002).
(I haven't actually read Getting Things Done, but I have friends who swear by it, some of whom talk to me about it, to which I tend to wind up responding with “oh, that's just locality of reference” or similar.)
Exhibit B: An implementation of hierarchical timing wheels for the fleshy-ape platform, due to David Allen (2002).
(I haven't actually read Getting Things Done, but I have friends who swear by it, some of whom talk to me about it, to which I tend to wind up responding with “oh, that's just locality of reference” or similar.)
2007-12-31
2007-12-30
On the importance of punctuation
Just now I was reading a slashdot article where a bunch of people where whining about some links or other being to some site called “myminicity”. Several mentions of that name later, it finally dawned on me that it was not, in fact, a compound of the usual English state-of-being suffix -ity and some hypothetical trendy nonsense word(s) I hadn't heard of yet (cf. “meme”); but, rather, was to be read as the words “my mini city”.
(I still don't know what the thing actually is, because I don't care.)
If only people had decided to use hyphens when gluing words together for domain names, then I wouldn't have to occasionally waste valuable seconds wondering WTF some string of half-pronounceable letters is supposed to be. (They also wouldn't run the risk of an unfortunate powergenitalia incident, but how often does that happen?) Ah well.
(I still don't know what the thing actually is, because I don't care.)
If only people had decided to use hyphens when gluing words together for domain names, then I wouldn't have to occasionally waste valuable seconds wondering WTF some string of half-pronounceable letters is supposed to be. (They also wouldn't run the risk of an unfortunate powergenitalia incident, but how often does that happen?) Ah well.
2007-12-27
The Sound of Silence
Let's say you have a video file of some sort with no audio stream, and a device that plays video files but can't handle ones with no audio. And you want to use the command-line tool ffmpeg, because you already have a script set up to use it to do whatever transcoding, but for the slight problem of sound.
Maybe you don't, but I did. Know that ffmpeg can take multiple input and/or output files and shuffle them around, in addition to decoding/encoding them. Know also that ffmpeg is documented in a manner both voluminous and not terribly approachable.
The wrong thing to do is to try to figure out how to extract the length of the video track and generate that many audio samples of silence. But, if one hasn't noticed the right part of the man page, this may be what one tries to do. (I did.)
The right thing to do is to use the -shortest flag, which directs ffmpeg to stop whenever any input stream reaches its end, rather than when all inputs are done. Some of you may be (but probably none of you are) thinking of various versions of the Scheme procedure map and their behavior on lists of unequal and/or infinite length, recently a point of contention in the discussion leading up to R6RS. And what is Unix if not a platform for stream processing? (Don't answer that.)
Thus: -f s8 -i /dev/zero -shortest, placed either before or after the regular input file, depending on whether input that does have audio should be silenced or not, respectively.
It's a not entirely inelegant solution to a mildly ridiculous problem.
Maybe you don't, but I did. Know that ffmpeg can take multiple input and/or output files and shuffle them around, in addition to decoding/encoding them. Know also that ffmpeg is documented in a manner both voluminous and not terribly approachable.
The wrong thing to do is to try to figure out how to extract the length of the video track and generate that many audio samples of silence. But, if one hasn't noticed the right part of the man page, this may be what one tries to do. (I did.)
The right thing to do is to use the -shortest flag, which directs ffmpeg to stop whenever any input stream reaches its end, rather than when all inputs are done. Some of you may be (but probably none of you are) thinking of various versions of the Scheme procedure map and their behavior on lists of unequal and/or infinite length, recently a point of contention in the discussion leading up to R6RS. And what is Unix if not a platform for stream processing? (Don't answer that.)
Thus: -f s8 -i /dev/zero -shortest, placed either before or after the regular input file, depending on whether input that does have audio should be silenced or not, respectively.
It's a not entirely inelegant solution to a mildly ridiculous problem.
2007-11-13
Quod Erat Covered-In-Bees
Trying to prove something in Coq is like operating a large, fly-by-wire bulldozer. On the one hand, you have to figure out what to shove where, or how to program the machine to do the shoving for you. On the other hand, you can use your superior human pattern-recognition skills to avoid driving into obstacles, and if you do get stuck, you have at least some chance of figuring out how to extricate yourself.
Trying to prove something in ACL2 is like imperiously proclaiming, “Robot servants, hear my command!” and then they run out into the junkyard picking up stuff and moving it around. And if you're lucky, and the task you've given them simple enough, they get the job done for you while you sit back on a lounge chair sipping a strawberry daiquiri. But if that's not the case, they run around doing exactly the wrong thing, until either they explode messily of their own accord or you call in an airstrike. And then you get to examine the flaming wreckage to figure out what went wrong, and what helpful hints you might be able to offer the robots next time.
Trying to prove something in ACL2 is like imperiously proclaiming, “Robot servants, hear my command!” and then they run out into the junkyard picking up stuff and moving it around. And if you're lucky, and the task you've given them simple enough, they get the job done for you while you sit back on a lounge chair sipping a strawberry daiquiri. But if that's not the case, they run around doing exactly the wrong thing, until either they explode messily of their own accord or you call in an airstrike. And then you get to examine the flaming wreckage to figure out what went wrong, and what helpful hints you might be able to offer the robots next time.
2007-11-09
This Morning's Compiler Meditation:
translating Knuth's “man or boy” test from ALGOL 60 into C, by reifying the activations as structs and closure-converting the call-by-name thunks. Like this.
I considered trying to run it by hand instead, but that would certainly take a lot of paper, and also as Knuth himself apparently failed in that task, I thought perhaps not.
I considered trying to run it by hand instead, but that would certainly take a lot of paper, and also as Knuth himself apparently failed in that task, I thought perhaps not.
2007-10-29
My Parity Iz Pastede On Yay
Sun's ZFS, to hear some people tell it, is the last filesystem anyone will ever need. It slices, it dices, and so on. Why, its magical powers are so great, it can give you all the advantages RAID-5 (or -6) without any of the historical drawbacks!
Except that it doesn't. Give you all of the benefits of RAID-[56], that is. In order for RAID-Z to do its thing, each filesystem block has to be divided into stripes independently. Which is great if you're dealing with a large file divided into 128kB blocks. And not so great if you're dealing with a single 5kB file that's been scattered across 10 data disks, one sector to each. Now you have to do I/O operations on all 10 disks to read the file back (and 12 to write it in the first place, assuming the use of double-parity raidz2), whereas with a conventional RAID with a more reasonable stripe size, it would take only one.
And that's the tradeoff. It might not sound like much, but it's a really big deal if you're trying to store mail in the Maildir format (one file per message!), or any other workload with a lot of small files and random access. System administrators dealing with such things may take to talking about “spindles” as an attribute of a disk array, as metonymy for the rate at which it can handle random I/Os. As in, the mail fileserver has been slow lately; looks like it needs more spindles. So you add disks to it, even though you have plenty of space free. Because it's all about how fast you can move the physical heads around.
In this light, what RAID-Z does is takes all of your disks and turns them into one spindle. Burning half the blocks to do mirroring instead doesn't sound so bad now, does it?
Except that it doesn't. Give you all of the benefits of RAID-[56], that is. In order for RAID-Z to do its thing, each filesystem block has to be divided into stripes independently. Which is great if you're dealing with a large file divided into 128kB blocks. And not so great if you're dealing with a single 5kB file that's been scattered across 10 data disks, one sector to each. Now you have to do I/O operations on all 10 disks to read the file back (and 12 to write it in the first place, assuming the use of double-parity raidz2), whereas with a conventional RAID with a more reasonable stripe size, it would take only one.
And that's the tradeoff. It might not sound like much, but it's a really big deal if you're trying to store mail in the Maildir format (one file per message!), or any other workload with a lot of small files and random access. System administrators dealing with such things may take to talking about “spindles” as an attribute of a disk array, as metonymy for the rate at which it can handle random I/Os. As in, the mail fileserver has been slow lately; looks like it needs more spindles. So you add disks to it, even though you have plenty of space free. Because it's all about how fast you can move the physical heads around.
In this light, what RAID-Z does is takes all of your disks and turns them into one spindle. Burning half the blocks to do mirroring instead doesn't sound so bad now, does it?
Subscribe to:
Posts (Atom)
